Running a Script
Hello, World
use v5.38; is the line to put at the top of every Perl script you write today: it turns on strict and warnings, and it enables say, which is print with the newline already on it.echo "Hello, World!"use v5.38;
say "Hello, World!";The semicolon is the visible change and the smallest one.
say corresponds to echo almost exactly, including that it adds the newline; print is the one that does not, which is the reverse of what shell programmers expect from printf. Without use v5.38 (or use strict; use warnings;) Perl behaves like a shell in the one way you do not want — undeclared variables spring into existence — so treat it as mandatory rather than optional.The shebang, and what replaces set -euo pipefail
The two headers are trying to do the same job — make the language stop being permissive — and Perl's is both shorter and more thorough, because it is checked when the script is compiled rather than while it runs.
#!/usr/bin/env bash
set -euo pipefail
# -e stop on a failed command
# -u an unset variable is an error
# -o pipefail a failure anywhere in a pipeline counts
echo "started"#!/usr/bin/env perl
use v5.38;
# use v5.38 implies:
# use strict an undeclared variable is a compile-time error
# use warnings suspicious operations say so
#
# There is no pipeline to leak a status, and a die is not ignorable.
say "started";use strict is the important half. A misspelled variable in Bash expands to the empty string and the script carries on; under strict, a misspelled variable in Perl is a compile error and the script does not start at all. That is a stronger guarantee than set -u, which only fires when the line is reached. The habit worth forming: never write a Perl script without that first line, exactly as you would never write a shell script without set -euo pipefail.Braces instead of do/done
qw(alpha beta) is a quoted word list — it saves the quotes and commas, and it is the closest thing Perl has to the shell's bare word list after in.for name in alpha beta; do
if [[ "$name" == "beta" ]]; then
echo "found $name"
fi
doneuse v5.38;
for my $name (qw(alpha beta)) {
if ($name eq 'beta') {
say "found $name";
}
}Braces replace the matched keyword pairs, and unlike C the braces are required even for a single statement, so the dangling-else class of bug does not exist. Note
eq rather than == for the string comparison: Perl has separate operators for numeric and string comparison (== against eq, < against lt), which is the same distinction Bash draws between -eq and == inside [[ ]] — with the spellings swapped, which catches everyone once.Variables & Sigils
Assignment, and the spaces that break it
my is the word to get into your fingers: it declares a variable in the enclosing block, which is what local does in a shell function — except that under use strict you cannot forget it, because omitting it will not compile.name="widget"
count=3
# name = "widget" ← this is a COMMAND called name, and it fails
echo "$name $count"use v5.38;
my $name = 'widget';
my $count = 3;
# Spaces are ordinary. The word my declares this variable in this scope.
say "$name $count";The dollar sign staying on the left-hand side is the thing shell programmers find strangest: in Bash you write
name= to set and $name to read, while in Perl the sigil is part of the variable's name and is always there. Once that clicks, the rest follows — and note that $count = 3 is a real number, not the one-character string that count=3 produces in the shell.The sigil says what you are asking for
🚨 The rule that makes Perl click: the sigil describes what you are getting, not what the container is.
@names is the array, and $names[1] is a scalar taken out of it — which is why it starts with a dollar even though the array does not.names=(alpha beta gamma)
echo "${names[1]}" # one element
echo "${names[@]}" # all of them
echo "${#names[@]}" # how many
declare -A role
role[web]=frontend
echo "${role[web]}"use v5.38;
my @names = qw(alpha beta gamma);
say $names[1]; # $ — one scalar out of the array
say "@names"; # @ — the whole array
say scalar @names; # how many
my %role = (web => 'frontend');
say $role{web}; # $ again — one scalar out of the hashThe shell has the same idea buried in worse syntax:
${names[1]} and ${names[@]} differ only by the subscript, and ${#names[@]} adds a hash on the front to mean something else entirely. Perl uses square brackets for arrays and curly braces for hashes, consistently, which is one fewer thing to look up. And a Perl hash key needs no quoting in the common case — $role{web} — where the shell's declare -A is easy to forget and silently ruins the subscript when you do.🚨 The quoting problem ends here too
The same demonstration as every shell-to-anything page on this site, because it is the same reason. The Bash loop runs twice for one filename; the Perl loop runs once.
file="my report.txt"
# Unquoted, the shell splits this into TWO words:
for word in $file; do
echo "got: [$word]"
done
echo "quoted: [$file]"use v5.38;
my $file = 'my report.txt';
# There is no splitting. A scalar is one value, everywhere.
for my $word ($file) {
say "got: [$word]";
}
say "quoted: [$file]";Perl does have a version of this hazard and it is worth knowing about rather than being surprised by: an array interpolated into a string (
"@names") is joined with spaces, and a list assigned to a scalar gives you its length rather than the list. But those are explicit conversions between explicit types, not an invisible re-splitting of every unquoted expansion, and neither is triggered by a space inside a filename.Arithmetic without a sub-language
Perl has one kind of number and works out whether it is an integer or a float as it goes, so
7 / 2 is 3.5 and no separate arithmetic mode is entered to find that out.count=7
echo "half: $(( count / 2 ))"
# Integers only. Anything with a decimal point leaves the shell:
echo "exact: $(echo "scale=3; 7 / 2" | bc)"use v5.38;
my $count = 7;
say "half: ", int($count / 2);
say "exact: ", $count / 2;
printf "rounded: %.3f\n", $count / 3;This alone retires
bc from a great many scripts. Note int() for the truncating division that $(( )) gives you — and note that int truncates toward zero rather than flooring, so int(-7/2) is -3, matching the shell. Perl also has the shell's numeric/string split in its operators rather than its syntax: + always adds numbers and . always joins strings, so "3" + 4 is 7 and 3 . 4 is "34", with no ambiguity about which was meant.Strings & Quoting
Single and double quotes mean the same things
The rule transfers unchanged, which is unusual on this page: double quotes interpolate variables, single quotes do not, and a backslash escapes the sigil.
name="widget"
echo "double: $name" # interpolates
echo 'single: $name' # literal
echo "escaped: \$name"use v5.38;
my $name = 'widget';
say "double: $name"; # interpolates
say 'single: $name'; # literal
say "escaped: \$name";Perl adds
qq{...} and q{...} as the same two quotes with a delimiter you choose, which is what saves you when the text is full of quotation marks — qq{he said "no"} needs no escaping at all. The habit that carries over badly is the shell's reflex to quote absolutely everything: in Perl, a variable is a variable whether or not it is inside quotes, so foo($name) is correct and foo("$name") is a needless copy.Substrings and suffixes
Perl's
substr takes an offset and a length, exactly as ${path:5:3} does — which is one of the few places a shell habit transfers with no adjustment at all.path="/var/log/system.log"
echo "${path##*/}" # basename
echo "${path%/*}" # dirname
echo "${path%.log}" # strip suffix
echo "${path:5:3}" # offset and lengthuse v5.38;
use File::Basename;
my $path = '/var/log/system.log';
say basename($path);
say dirname($path);
say ($path =~ s/\.log$//r); # /r returns a copy, leaving $path alone
say substr($path, 5, 3); # offset and length — same as the shellThe
/r flag on the substitution is worth learning early: without it, s/// modifies the variable in place and returns how many replacements it made, which is a surprise if you expected the new string. File::Basename is core, so it needs no installation. For anything involving real paths prefer it — or File::Spec — over pattern-matching the string, for the same reason you would in any language.Joining and splitting
split and join are ordinary functions taking a pattern and a list. The shell has to reassign IFS — a global that changes how the whole shell parses words — and then put it back.line="alpha,beta,gamma"
IFS=',' read -r -a parts <<< "$line"
echo "count: ${#parts[@]}"
echo "first: ${parts[0]}"
joined="$(IFS='|'; echo "${parts[*]}")"
echo "$joined"use v5.38;
my $line = 'alpha,beta,gamma';
my @parts = split /,/, $line;
say "count: ", scalar @parts;
say "first: $parts[0]";
say join '|', @parts;The
IFS dance in the left column is the source of a specific and nasty bug: forget the subshell parentheses around the assignment and IFS stays changed for the rest of the script, so unrelated code splits its words on the wrong character. Perl's split takes the separator as an argument, so nothing global changes, and the separator is a full regular expression — split /\s*,\s*/ handles the spaces around the commas in the same call.Arrays & Hashes
Arrays
say "- $_" for @names; is a statement modifier — the loop written after the thing it repeats, with $_ as the current element. It is idiomatic Perl and reads well for one-line bodies.names=(alpha beta gamma)
names+=(delta)
echo "count: ${#names[@]}"
for name in "${names[@]}"; do
echo "- $name"
done
echo "sorted: $(printf '%s\n' "${names[@]}" | sort -r | tr '\n' ' ')"use v5.38;
my @names = qw(alpha beta gamma);
push @names, 'delta';
say "count: ", scalar @names;
say "- $_" for @names;
say "sorted: @{[ reverse sort @names ]}";Sorting is the row to notice. In the shell, sorting an array means printing it, piping it through
sort, and reading the text back — a subprocess and a round trip through a stream. In Perl sort is an operator on the list itself, it takes a comparison block for anything non-alphabetical (sort { $a <=> $b } @numbers for numeric order), and nothing leaves the process. The @{[ ... ]} around it is the trick for running an expression inside a string.Hashes
A hash is written as a list of key-value pairs, and the fat comma
=> is a comma that also quotes the word on its left — so apple => 2 needs no quotation marks.declare -A counts
counts[apple]=2
counts[pear]=5
echo "apple: ${counts[apple]}"
for key in "${!counts[@]}"; do
echo "$key = ${counts[$key]}"
done | sortuse v5.38;
my %counts = (apple => 2, pear => 5);
say "apple: $counts{apple}";
for my $key (sort keys %counts) {
say "$key = $counts{$key}";
}Hashes are where Perl earns its reputation and where the shell most obviously runs out.
keys, values, exists and delete are operators; a hash may be passed to a subroutine, returned from one, nested inside another hash, and sorted by value with a comparison block. Bash associative arrays can do none of those, need version 4 or later, and hold only strings — and the ${!counts[@]} spelling for the keys is a piece of punctuation nobody remembers unaided.Counting things
$count{$_}++ is the whole counting algorithm. The key springs into existence with the value zero the first time it is incremented, which is exactly what makes this idiom two words long.printf '%s\n' banana apple cherry apple |
sort | uniq -c | sort -rn |
awk '{ print $2 ": " $1 }'use v5.38;
my @fruit = qw(banana apple cherry apple);
my %count;
$count{$_}++ for @fruit;
for my $name (sort { $count{$b} <=> $count{$a} || $a cmp $b } keys %count) {
say "$name: $count{$name}";
}This is the single most-typed line in the history of Perl, and it is the reason so many
sort | uniq -c | sort -rn pipelines became Perl scripts. The sort block shows the other half of the gain: <=> compares numbers, cmp compares strings, and || chains them into a tiebreak — sort by count descending, then by name — which in the shell means either a second sort -k with the right field numbers or giving up.Conditionals & Loops
if, and what counts as true
Truth in Perl is a property of the value, as in most languages:
0, "", "0" and undef are false and everything else is true. Truth in the shell is an exit status, where zero means success — the opposite convention, which is the thing to keep straight.count=0
name=""
if [[ "$count" -eq 0 ]]; then echo "count is zero"; fi
if [[ -z "$name" ]]; then echo "name is empty"; fi
if [[ "$name" != "widget" ]]; then echo "not the widget"; fiuse v5.38;
my $count = 0;
my $name = '';
say "count is zero" if $count == 0;
say "name is empty" if !length $name;
say "not the widget" if $name ne 'widget';🚨
"0" being false is the Perl-specific trap, and it has no shell equivalent: a string containing just the digit zero is false, so if ($answer) on user input that is legitimately "0" takes the wrong branch. Use if (length $answer) or if (defined $answer) when the empty case is what you actually mean. Note also the operator pairs — ==/eq, !=/ne, </lt — which are the same numeric-versus-string distinction [[ ]] makes with -eq and ==, spelled the other way round.Loops
1 .. 3 is the range operator and it includes both ends, exactly as {1..3} does — so unlike most languages on this site, the boundary does not need rethinking.for index in {1..3}; do
echo "index $index"
done
count=0
while (( count < 2 )); do
echo "while $count"
count=$(( count + 1 ))
doneuse v5.38;
for my $index (1 .. 3) {
say "index $index";
}
my $count = 0;
while ($count < 2) {
say "while $count";
$count++;
}One real gain:
1 .. $limit works, where the shell's {1..$limit} does not. Brace expansion happens before variables are substituted, so the shell version produces the literal text {1..3} and the loop runs once over a nonsense value — a bug that has caught everyone at least once and that seq exists to work around.case becomes a dispatch table
Perl has no
case statement worth using, and the idiomatic replacement is better than one: a hash whose values are anonymous subroutines, looked up by key. // is the defined-or operator, supplying the fallback.for value in start stop other; do
case "$value" in
start) echo "$value: starting" ;;
stop) echo "$value: stopping" ;;
*) echo "$value: unknown" ;;
esac
doneuse v5.38;
my %action = (
start => sub { say "$_[0]: starting" },
stop => sub { say "$_[0]: stopping" },
);
for my $value (qw(start stop other)) {
my $handler = $action{$value} // sub { say "$_[0]: unknown" };
$handler->($value);
}The dispatch table is a genuinely different capability rather than a translation. The handlers are values, so they can be added at run time, built from a configuration file, wrapped for logging, or tested one at a time — none of which a
case statement allows. For the simple string-matching cases an if/elsif chain is perfectly good and clearer; reach for the table when there are more than about four branches or when the set of branches is not fixed.Regular Expressions
Regular expressions stop being an argument
The match returns its capture groups as a list, so they can be assigned straight into named variables. That is the whole difference: in Bash they land in a global array called
BASH_REMATCH and you index it by number.line="ERROR 2026-09-09 disk full"
if [[ "$line" =~ ^([A-Z]+)\ ([0-9-]+)\ (.*)$ ]]; then
echo "level: ${BASH_REMATCH[1]}"
echo "date: ${BASH_REMATCH[2]}"
echo "text: ${BASH_REMATCH[3]}"
fiuse v5.38;
my $line = 'ERROR 2026-09-09 disk full';
if (my ($level, $date, $text) = $line =~ /^(\w+) (\S+) (.*)$/) {
say "level: $level";
say "date: $date";
say "text: $text";
}Three things improve at once. The captures have names, so line 40 of the script is readable. The result is a value rather than a global, so an unrelated match two lines later cannot clobber it. And the pattern is not passing through the shell's quoting first — notice the backslash-escaped spaces the Bash version needs, and the fact that quoting the pattern in
[[ ]] would turn the whole regular expression off. Perl was designed around this operator and it shows.sed s/// becomes s///
The syntax is the same syntax —
sed borrowed it first and Perl extended it. s{...}{...} lets you pick the delimiter, which is how the slashes in the replacement avoid being escaped.echo "hello world" | sed 's/o/0/g'
echo "2026-09-09" | sed -E 's/([0-9]{4})-([0-9]{2})-([0-9]{2})/\3\/\2\/\1/'use v5.38;
my $greeting = 'hello world';
(my $leet = $greeting) =~ s/o/0/g;
say $leet;
my $date = '2026-09-09';
say $date =~ s{(\d{4})-(\d{2})-(\d{2})}{$3/$2/$1}r;Two Perl-side gains. The replacement is Perl code if you ask for it (
s/(\d+)/$1 * 2/e doubles every number, which sed cannot do at all), and the captures are $1, $2, $3 rather than \1 after shell quoting has had a turn. The /r flag returns the modified copy and leaves the original alone; without it, s/// modifies the variable in place and returns a count, which is the single most common surprise for someone arriving from sed.Finding every match
A match with
/g in list context returns every match as a list. They are numbers, so the sum is a function call, where the shell has to strip the non-digits from each word by hand and keep a running total.text="a1 b22 c333"
# Extracting every number means stripping the non-digits from each word
# by hand, and then adding them up in a second pass:
total=0
for word in $text; do
digits="${word//[^0-9]/}"
echo "$digits"
total=$(( total + digits ))
done
echo "sum: $total"use v5.38;
use List::Util qw(sum0);
my $text = 'a1 b22 c333';
my @numbers = $text =~ /(\d+)/g;
say for @numbers;
say "sum: ", sum0(@numbers);The shell has no way to say "give me every match" —
[[ =~ ]] finds one, BASH_REMATCH holds that one, and getting the rest means either a loop that eats the string as it goes or shelling out to grep -o. Perl returns them as a list. List::Util is core — no installation — and brings sum0, max, min, first, any and uniq, which between them retire a large fraction of the awk in a typical script.sed, awk and cut, Collapsed
awk fields become split
split ' ' with a literal single-space string is a special case in Perl: it splits on any run of whitespace and discards leading whitespace, which is exactly what awk does by default.printf '%s\n' "web1 frontend 8080" "db1 database 5432" |
awk '{ print $2 " on " $1 ":" $3 }'use v5.38;
for my $line ('web1 frontend 8080', 'db1 database 5432') {
my ($host, $role, $port) = split ' ', $line;
say "$role on $host:$port";
}awk is very good at this and a one-line field extraction should stay one line of awk. What Perl offers is names instead of
$2, and the fact that the same program can then do something the pipeline cannot — look a value up in a hash, keep a running total in a real number, or write two output files at once. The line where a script stops being a pipeline and starts being a program is usually the line where a second stage needs to remember something.Four stages become one program
grep and map are Perl operators taking a block, and they compose right to left in the same order the pipeline runs left to right: filter, then transform, then sort.printf '%s\n' "INFO ok" "ERROR disk full" "INFO fine" "ERROR timeout" |
grep '^ERROR' |
sed 's/^ERROR //' |
sort |
awk '{ print NR ". " $0 }'use v5.38;
my @lines = ('INFO ok', 'ERROR disk full', 'INFO fine', 'ERROR timeout');
my @problems = sort map { s/^ERROR //r } grep { /^ERROR/ } @lines;
my $number = 0;
say ++$number, ". $_" for @problems;This is the argument of the whole page in one row. The left column is a small program written in four languages — shell quoting, a basic regular expression, an extended one, and awk — each with its own escaping rules, running as four processes. The right column is one language, one set of rules, one process, and every intermediate value is available to be counted, reused or written somewhere else. The right column is also, honestly, harder to read at a glance than the pipeline; the trade is that it stays readable as the job grows, and the pipeline does not.
sed -i becomes perl -i
Perl invented
-i and sed is the one people know it from. As a one-liner the two are the same length; the reason to write it out longhand is that a script you can test beats a string you cannot.printf '%s\n' "level=info" "path=/tmp" > /tmp/settings.conf
sed -i.backup 's/^level=.*/level=debug/' /tmp/settings.conf
cat /tmp/settings.confuse v5.38;
# The one-liner form, which is what you would actually type:
# perl -i.backup -pe 's/^level=.*/level=debug/' settings.conf
#
# The same thing inside a script, so it can be tested and reused:
my $path = '/tmp/settings.conf';
open my $write, '>', $path or die "cannot write: $!";
print {$write} "level=info\npath=/tmp\n";
close $write;
open my $read, '<', $path or die "cannot read: $!";
my @lines = <$read>;
close $read;
s/^level=.*/level=debug/ for @lines;
open $write, '>', $path or die "cannot write: $!";
print {$write} @lines;
close $write;
print for @lines;The
for @lines at the end of the substitution is doing something the shell cannot: $_ is an alias to each element, so modifying it modifies the array in place. That is why the loop needs no assignment. If the edit has any conditional logic in it — change this line only when that other line said something — the one-liner stops being enough, and the script form is where you were always going to end up.Files & Directories
Reading a file
open … or die "…: $!" is the idiom, and the $! holds the system error message. The shell's redirection is shorter; what it does not do is tell you why it failed.printf '%s\n' alpha beta gamma > /tmp/example.txt
while IFS= read -r line; do
echo "[$line]"
done < /tmp/example.txt
echo "lines: $(wc -l < /tmp/example.txt | tr -d '[:space:]')"use v5.38;
open my $write, '>', '/tmp/example.txt' or die "cannot write: $!";
say {$write} $_ for qw(alpha beta gamma);
close $write;
open my $read, '<', '/tmp/example.txt' or die "cannot read: $!";
my $count = 0;
while (my $line = <$read>) {
chomp $line;
say "[$line]";
$count++;
}
close $read;
say "lines: $count";chomp is the one to remember: <$read> keeps the newline on the end of each line, where the shell's read strips it. Forgetting chomp produces output with blank lines in it and comparisons that mysteriously fail, and it is the most common first-week Perl bug. The three-argument open shown here is the only form to use — the older two-argument version interprets characters in the filename and is a genuine security hazard.File tests are the same letters
Perl borrowed the file test operators from the shell letter for letter —
-e, -f, -d, -s, -r, -w, -x all mean what they mean in [[ ]].printf 'x' > /tmp/probe.txt
[[ -e /tmp/probe.txt ]] && echo "exists"
[[ -f /tmp/probe.txt ]] && echo "regular file"
[[ -d /tmp ]] && echo "directory"
[[ -s /tmp/probe.txt ]] && echo "not empty"use v5.38;
open my $handle, '>', '/tmp/probe.txt' or die "cannot write: $!";
print {$handle} 'x';
close $handle;
say "exists" if -e '/tmp/probe.txt';
say "regular file" if -f '/tmp/probe.txt';
say "directory" if -d '/tmp';
say "not empty" if -s '/tmp/probe.txt';One difference is useful rather than annoying:
-s returns the file's size rather than merely a true value, so it can be used in arithmetic. Another is -M and -A, giving age in days since the script started, which the shell has no equivalent for at all. And because these are expressions rather than commands, they combine with ordinary boolean operators instead of needing [[ ... && ... ]].Globbing
glob takes the same pattern the shell would and returns a list of matching paths. The important difference is what happens when nothing matches.mkdir -p /tmp/globdemo
: > /tmp/globdemo/one.txt
: > /tmp/globdemo/two.txt
: > /tmp/globdemo/three.log
for file in /tmp/globdemo/*.txt; do
echo "found: $(basename "$file")"
doneuse v5.38;
use File::Basename;
mkdir '/tmp/globdemo' unless -d '/tmp/globdemo';
for my $name (qw(one.txt two.txt three.log)) {
open my $handle, '>', "/tmp/globdemo/$name" or die $!;
close $handle;
}
for my $file (sort glob '/tmp/globdemo/*.txt') {
say "found: ", basename($file);
}An unmatched shell glob expands to the pattern itself, so a loop body runs once with a filename that does not exist — the bug
shopt -s nullglob exists to prevent and that nobody remembers to enable. glob returns an empty list, so the loop simply does not run. For walking a tree, File::Find is core and replaces find without a subprocess or a -print0/xargs -0 dance over filenames with spaces in them.Running Other Programs
Capturing a command's output
qx{...} is Perl's command substitution — the same thing backticks do, written so the delimiter is visible. The whole string goes to a shell, so shell quoting rules apply inside it.output="$(printf '%s\n' gamma alpha beta | sort)"
echo "$output"
echo "status: $?"use v5.38;
my $output = qx{printf '%s\\n' gamma alpha beta | sort};
print $output;
say "status: ", $? >> 8;🚨
$? holds the wait status rather than the exit code, so it must be shifted right by eight to get the number the shell puts in $?. That is a genuine wart. The bigger warning is that qx and the one-argument system both hand your string to a shell, so a filename with a space or a semicolon in it is an injection: use the list form, system('sort', $file), whenever any part of the command comes from data.Reading from a program as a file
The
-| mode opens a program's standard output as a file handle, and passing the arguments as a list means no shell is involved — so nothing in the data can be interpreted as a command.printf '%s\n' 3 1 4 1 5 | sort -n | while IFS= read -r value; do
echo "got $value"
doneuse v5.38;
# A pipe opened as a file handle: the program's output is read line by line.
# The program is $^X, the perl running this script, so it exists everywhere.
open my $numbers, '-|', $^X, '-e', 'print "$_\n" for 3, 1, 4, 1, 5'
or die "cannot run: $!";
my @values = sort { $a <=> $b } map { chomp; $_ } <$numbers>;
close $numbers;
say "got $_" for @values;Notice that the sorting moved out of the pipeline and into Perl, which is the usual outcome: once the data is inside the program there is rarely a reason to send it back out to a utility. The list form of
open is the safe one and is worth making a habit, exactly as it is in every other language on this site. There is a matching |- mode for writing to a program's standard input.The environment
%ENV is an ordinary hash that happens to be wired to the process environment, so everything you know about hashes applies — exists, delete, iteration over keys.export REPORT_LEVEL=verbose
echo "level: $REPORT_LEVEL"
echo "set? $([[ -n "${REPORT_LEVEL:-}" ]] && echo yes || echo no)"
echo "unset reads as: [${NOT_SET_ANYWHERE:-}]"use v5.38;
$ENV{REPORT_LEVEL} = 'verbose';
say "level: $ENV{REPORT_LEVEL}";
say "set? ", ($ENV{REPORT_LEVEL} ? 'yes' : 'no');
say "unset reads as: [", ($ENV{NOT_SET_ANYWHERE} // ''), "]";Assigning to
%ENV is the equivalent of export: it affects processes this script starts afterwards and nothing else. The shell's distinction between a plain variable and an exported one has no counterpart, because a Perl variable is not in the environment at all unless you put it there — which is the clearer arrangement of the two.Subroutines
Parameters get names
Perl gained real named parameters with defaults in 5.36, and
use v5.38 switches them on. Older code unpacks @_ by hand — my ($name, $greeting) = @_; — which you will still meet everywhere.greet() {
local name="$1"
local greeting="${2:-Hello}"
echo "$greeting, $name!"
}
greet "world"
greet "world" "Goodbye"use v5.38;
sub greet ($name, $greeting = 'Hello') {
say "$greeting, $name!";
}
greet('world');
greet('world', 'Goodbye');The
local that the shell version depends on is not needed, because my in a subroutine is already scoped to it and use strict will not let you forget. Perl's own local is a different and rarer thing — a temporary value for a global, restored when the block exits — and it is not the keyword you want here.🚨 return returns a value
This is the row that explains most of what looks odd in large shell scripts. A shell function cannot hand a value back, so it prints one, and the caller forks a subshell to catch it.
# A shell function returns an EXIT STATUS: 0-255 and nothing else.
double() {
echo $(( $1 * 2 )) # the "value" is printed
}
result="$(double 21)" # and captured by forking a subshell
echo "result: $result"
# Returning two things means printing them and splitting them back up.
min_max() { echo "1 9"; }
read -r low high <<< "$(min_max)"
echo "low=$low high=$high"use v5.38;
sub double ($value) { return $value * 2 }
my $result = double(21);
say "result: $result";
# A list comes back as a list — no printing, no parsing, no subshell.
sub min_max (@values) {
my @sorted = sort { $a <=> $b } @values;
return ($sorted[0], $sorted[-1]);
}
my ($low, $high) = min_max(4, 9, 1);
say "low=$low high=$high";Three costs vanish at once. There is no fork, so a script can be decomposed into small subroutines without paying for it. The value keeps its type, so a number stays a number and a list stays a list. And returning two things is
return ($low, $high) rather than an encoding the caller has to reverse — which is the point at which shell scripts usually start passing data through global variables and become hard to follow.References & Real Structure
The structure the shell cannot hold
A hash whose values are hashes. The curly braces inside the list create anonymous hashes and what is stored is a reference to each — which is the one genuinely new idea Perl asks a shell programmer to learn.
# No nesting. The workaround is a flattened key that you parse back
# out, and it breaks on any value containing your separator.
declare -A machine
machine["web1:role"]=frontend
machine["web1:port"]=8080
machine["db1:role"]=database
for key in "${!machine[@]}"; do
host="${key%%:*}"; field="${key##*:}"
echo "$host $field = ${machine[$key]}"
done | sortuse v5.38;
my %machine = (
web1 => { role => 'frontend', port => 8080 },
db1 => { role => 'database' },
);
for my $host (sort keys %machine) {
my $fields = $machine{$host};
for my $field (sort keys %$fields) {
say "$host $field = $fields->{$field}";
}
}References are how Perl holds anything nested:
\@array and \%hash make one, $reference->{key} and $reference->[0] follow one, and %$reference treats it as a whole hash again. It takes an afternoon and it is the whole difference between a script that can describe its data and one that has to encode it into a string. Note the port is a real number here, comparable and addable, where every value in the left column is text.Passing a collection to a function
The backslash makes a reference, and the subroutine receives one scalar that points at the caller's array. It is the same idea as passing a pointer, and Perl's references were deliberately modeled on C's pointers.
# An array cannot be passed. The workarounds are all bad:
# - pass the NAME and use nameref (bash 4.3+, and cryptic)
# - flatten to a string and split it back
# - use a global and hope
process() {
local -n reference="$1" # nameref: pass the variable's NAME
echo "count: ${#reference[@]}"
echo "first: ${reference[0]}"
}
names=(alpha beta gamma)
process namesuse v5.38;
sub process ($names) {
say "count: ", scalar @$names;
say "first: $names->[0]";
}
my @names = qw(alpha beta gamma);
process(\@names); # \ makes a reference to the arrayBash 4.3 added namerefs, which pass the variable's name and let the callee alias it — it works, it is not widely known, and it puts the caller's variable name into the callee's hands, which is exactly backwards. Perl's version has the ordinary property you expect: the reference is a value, so it can be stored in another structure, returned, or put in a list of references. Note that a plain
@names passed to a subroutine is flattened into the argument list, which is why the reference is needed to keep it whole.Structured output
Look at the first line of the shell output: the embedded quotation marks break the JSON, silently, and only for data that contains them.
JSON::PP is core, so it needs no installation.name='say "hello"'
count=3
# Concatenating by hand breaks the moment a value has a quote in it:
echo "{\"name\": \"$name\", \"count\": $count}"
# jq -n is the correct shell answer:
jq -nc --arg name "$name" --argjson count "$count" \
'{name: $name, count: $count}'use v5.38;
use JSON::PP;
my $data = {
name => 'say "hello"',
count => 3,
tags => ['alpha', 'beta'],
};
say JSON::PP->new->canonical->encode($data);Encoding takes a data structure rather than text, so quoting is not a decision the caller can get wrong — the same argument that applies to building SQL, HTML or a shell command.
canonical sorts the keys so the output is reproducible, which matters when it is compared or checked into version control. Decoding with decode_json gives you nested hashes and arrays, where jq can only hand back text for the shell to re-parse.Errors & Exit Codes
set -e becomes die
die throws, and eval { } catches — it is try/catch with older spelling, and the error lands in $@. The trailing 1 inside the block is the idiom that makes the or fire only on failure.set -euo pipefail
if ! false; then
echo "the command failed, and we noticed"
fi
echo "still running"use v5.38;
sub might_fail () { die "the command failed\n" }
eval { might_fail(); 1 } or do {
my $error = $@;
chomp $error;
say "$error, and we noticed";
};
say "still running";Perl 5.34 added real
try/catch blocks under use feature 'try', and they are worth using in new code — the eval form shown here is what you will read in everything older. Either way, the gain over set -e is that a failure carries a message and a type rather than a number from 0 to 255, and that the rule for what propagates is one sentence rather than the list of exceptions set -e has accumulated.Exiting with a status
Both write the diagnostic to standard error —
>&2 on the left, STDERR on the right — so a caller can separate it from the result.check() {
if [[ "$1" == "bad" ]]; then
echo "fatal: bad input" >&2
return 2
fi
echo "ok"
}
check good
check bad || echo "caller saw status $?"use v5.38;
sub check ($value) {
die "bad input\n" if $value eq 'bad';
say 'ok';
}
check('good');
eval { check('bad'); 1 } or do {
my $error = $@; chomp $error;
print STDERR "fatal: $error\n";
say "caller saw: $error";
# exit 2 here would hand that status back to the shell.
};An exit status is one number out of 256 and the message is unstructured text on another stream, so a caller can only tell failures apart by convention. A thrown error carries whatever you put in it, including an object with fields, and
die automatically appends the file and line when the message does not end in a newline — which is why the examples here end theirs with \n, to suppress it. exit still exists for the boundary where a Perl script really does have to answer a shell.One-Liners & Arguments
perl -ne is a pipeline stage
Perl is a perfectly good pipeline stage and the flags are worth memorizing:
-n wraps your program in a loop over input lines, -a splits each line into @F, and -l deals with the newlines. That combination is awk, with Perl's regular expressions.printf '%s\n' "alpha 3" "beta 14" "gamma 15" |
awk '$2 > 10 { print $1 }'use v5.38;
# As a one-liner, which is how this is actually used:
# ... | perl -lane 'print $F[0] if $F[1] > 10'
#
# -e the program -n loop over input lines
# -l handle newlines -a autosplit each line into @F
#
# The same thing as a script:
for my $line ('alpha 3', 'beta 14', 'gamma 15') {
my @fields = split ' ', $line;
say $fields[0] if $fields[1] > 10;
}So moving to Perl does not mean abandoning the shell — a Perl one-liner sits in a pipeline exactly where awk did, and reaches for a hash or a module when the job outgrows a single expression.
-p is -n plus printing each line, which with -i is the in-place editor from the Text section. The migration this page describes is usually gradual for exactly this reason: one stage at a time.getopts becomes Getopt::Long
Each option is declared once with both spellings, and
=s says it takes a string. Getopt::Long is core, so this needs nothing installed. bash script.sh -v -o report.txt
verbose=0
output=""
while getopts "vo:" option; do
case "$option" in
v) verbose=1 ;;
o) output="$OPTARG" ;;
*) echo "usage: report.sh [-v] [-o FILE]" >&2; exit 1 ;;
esac
done
echo "verbose=$verbose output=${output:-none}"
# Short options only. --output means writing the loop yourself. perl main.pl --verbose --output report.txt
use v5.38;
use Getopt::Long;
my $verbose = 0;
my $output;
GetOptions(
'verbose|v' => \$verbose,
'output|o=s' => \$output,
) or die "usage: report.pl [--verbose] [--output FILE]\n";
say "verbose=$verbose output=", ($output // 'none');Long options are the whole point.
getopts handles single letters and nothing else, so every shell script that accepts --dry-run contains a hand-rolled parsing loop with its own bugs and its own usage message that drifts out of date. Getopt::Long also does abbreviation, negation (--no-verbose), repeated options collecting into an array, and =i for values that must be integers.Modules & Growing Up
source becomes use
This is the difference that decides whether a script can grow past one file.
source is textual inclusion into one namespace; use loads a package and imports only the names you ask for.# One shared global namespace. Two helper files that both define
# "log" silently overwrite each other, and which one wins depends
# on the order of the source lines.
# source ./helpers.sh
log() { echo "[log] $*"; }
log "started"use v5.38;
# Each module is its own package with its own namespace:
# use Helpers qw(log_line); # imports just what you name
# Helpers::log_line('started'); # or call it fully qualified
#
# Two modules may both define log_line without disturbing each other.
sub log_line (@parts) { say '[log] ', join ' ', @parts }
log_line('started');A module is loaded once however many files ask for it, so there is no include-guard problem, and the
use lines at the top of a file are an honest inventory of what it depends on. The other half is CPAN, which is the largest and oldest module archive of its kind — and the counterweight is that a script depending on CPAN modules needs them installed wherever it runs, which is precisely the deployment simplicity that keeps things in shell.Testing
Test::More is core and emits TAP, the Test Anything Protocol — which Perl invented and which most other languages now have a consumer for. Nothing is installed to run this.# No assertion and no runner in the box. Either a third-party
# framework (bats, shunit2) or this:
assert_equal() {
if [[ "$1" != "$2" ]]; then
echo "FAIL: expected [$2], got [$1]" >&2
return 1
fi
echo "ok"
}
double() { echo $(( $1 * 2 )); }
assert_equal "$(double 21)" "42"use v5.38;
use Test::More;
sub double ($value) { return $value * 2 }
is(double(21), 42, 'doubles a number');
is(double(0), 0, 'handles zero');
done_testing();Testability is the quiet reason scripts get rewritten. A shell function can only be tested through the text it prints, so the test asserts on formatting as well as behavior and breaks when a message is reworded. A subroutine returns a value, so the test checks the value and ignores the presentation — and
prove, also core, runs a directory of such files and summarizes them.When to stay in the shell
A page like this one can read as an argument that shell scripts are bad. They are not, and the last row should say so plainly: the left column is exactly right as it stands.
#!/usr/bin/env bash
# This should NOT be rewritten. It is glue: every line runs a
# program, which is what a shell is for.
set -euo pipefail
mkdir -p /tmp/release
printf '%s\n' one two three > /tmp/release/manifest.txt
sort /tmp/release/manifest.txt | tee /tmp/release/sorted.txt | wc -luse v5.38;
# The same job in Perl: longer, no clearer, and it gives up the thing
# the shell is best at. Rewrite when the script has grown LOGIC —
# nested data, arithmetic, error handling that distinguishes failures,
# more than a few options — not merely because it has grown long.
mkdir '/tmp/release' unless -d '/tmp/release';
open my $write, '>', '/tmp/release/manifest.txt' or die $!;
say {$write} $_ for qw(one two three);
close $write;
open my $read, '<', '/tmp/release/manifest.txt' or die $!;
my @lines = sort <$read>;
close $read;
open $write, '>', '/tmp/release/sorted.txt' or die $!;
print {$write} @lines;
close $write;
say scalar @lines;The test is not length, it is whether the script has logic — data with more than one level, arithmetic beyond counting, error handling that must tell failures apart, several options, or anything the next person will have to modify under pressure. Orchestration and glue belong in the shell. And the migration need not be all at once: a Perl one-liner is a pipeline stage, so the usual path is to replace the gnarliest stage first and leave the rest alone.