PONYλM2Modula-2

Bash.CodeCompared.To/Ruby

An interactive executable cheatsheet comparing Bash and Ruby

Bash 5.3 Ruby 4.0
Running a Script
Hello, World
puts is echo: it writes its argument and adds the newline. The only visible change is the word.
echo "Hello, World!"
puts "Hello, World!"
Two neighbors worth meeting now. print is the one that does not add a newline, which is the reverse of the habit printf gives you. And p prints a debugging representation — p "a" shows the quotation marks — which is the fastest way to see whether a value is the string "3" or the number 3, a distinction the shell never made you think about.
The shebang, and what replaces set -euo pipefail
The three options on the left are the shell trying to behave the way Ruby already does by default. There is no Ruby equivalent to write, because there is nothing permissive to turn off.
#!/usr/bin/env bash set -euo pipefail # -e stop on a failed command # -u an unset variable is an error # -o pipefail a failure anywhere in a pipeline counts echo "started"
#!/usr/bin/env ruby # There is nothing to switch on. An error raises and stops the script, # an undefined local variable raises NameError, and there is no pipeline # to swallow a failure in the middle. puts "started"
The one thing worth adding at the top of a Ruby script is # frozen_string_literal: true, which is about performance and mutation rather than about safety. Note the asymmetry that catches shell programmers: an undefined local variable raises, but an undefined hash key quietly returns nil — so settings[:levl] is the typo that survives, and settings.fetch(:levl) is how you make it raise.
end instead of fi and done
%w[alpha beta] is a word array — the shell's bare list after in, with the same "no quotes needed" convenience. The |name| between the pipes is the loop variable, and it belongs to the block rather than to the surrounding script.
for name in alpha beta; do if [[ "$name" == "beta" ]]; then echo "found $name" fi done
%w[alpha beta].each do |name| if name == "beta" puts "found #{name}" end end
Ruby closes every block with end rather than with a keyword matched to the opener, so there is no fi/done/esac to keep straight — but there is also no visual reminder of what is being closed, which is why indentation matters more in practice than the parser cares. The interpolation #{name} works only in double quotes, exactly as $name does in the shell.
Variables & Quoting
Assignment, and the spaces that break it
The space around the equals sign is the first thing to unlearn, and the type is the second: count=3 in the shell is a one-character string, while count = 3 in Ruby is an integer.
name="widget" count=3 # name = "widget" ← this is a COMMAND called name, and it fails echo "$name $count"
name = "widget" count = 3 # Spaces are ordinary. count is the NUMBER three, not the string "3". puts "#{name} #{count}"
That type distinction is the one to hold on to, because Ruby will not convert for you: "3" + 4 raises a TypeError rather than quietly producing 7 or "34". After years of a language where everything is text and conversion is invisible, this reads as strictness — and it is the same strictness that makes count + 1 mean addition without a $(( )) around it.
🚨 The quoting problem ends
The single best reason to move a script off the shell, and it deserves to be seen rather than described: the Bash loop runs twice for one filename.
file="my report.txt" # Unquoted, the shell splits this into TWO words and the loop runs twice: for word in $file; do echo "got: [$word]" done echo "quoted: [$file]"
file = "my report.txt" # There is no splitting. A string is one value, everywhere, always. [file].each do |word| puts "got: [#{word}]" end puts "quoted: [#{file}]"
Word splitting is not a bug, it is the 1978 design — an unquoted expansion is re-split on whitespace and then glob-expanded, because that made the shell convenient to type at. Every "$variable" you have written is you switching it off. In Ruby a string is a value like any other: passing it, storing it, printing it and handing it to a method never splits it, and the whole category of filename-with-a-space bugs stops existing rather than being defended against.
Default values
|| and ||= line up with :- and := almost exactly, and they read better because they are ordinary operators rather than punctuation inside an expansion.
unset LEVEL echo "${LEVEL:-info}" # use a default, do not assign settings="" : "${settings:=warn}" # assign the default if unset or empty echo "settings is now: $settings"
level = nil puts level || "info" # || falls back on nil or false settings = nil settings ||= "warn" # assign only if nil or false puts "settings is now: #{settings}"
🚨 The one difference that bites: the shell treats the empty string as missing, and Ruby does not. "" is truthy in Ruby — only nil and false are falsy — so name || "default" keeps an empty string where ${name:-default} would replace it. When you mean "empty counts as missing", write name.to_s.empty? ? "default" : name, and be glad the decision is visible.
Arithmetic without a sub-language
No $(( )) and no bc. Arithmetic is just arithmetic, and whether you get an integer or a float depends on the operands rather than on which syntax you wrapped the expression in.
count=7 echo "half: $(( count / 2 ))" # Integers only — anything with a decimal point leaves the shell: echo "exact: $(echo "scale=3; 7 / 2" | bc)"
count = 7 puts "half: #{count / 2}" # integer division, both are integers puts "exact: #{count / 2.0}" # one float makes the whole thing float puts "rounded: #{(count / 3.0).round(3)}"
Ruby integers are arbitrary precision, so 2 ** 200 is exact rather than overflowing. The trap in the other direction is that 7 / 2 is 3, not 3.5 — integer division, the same answer $(( )) gives — so a stray integer silently discards your fraction. Beyond floats there is BigDecimal for money and Rational for exact ratios (1r/3), neither of which needs a subprocess.
Strings
Interpolation and quoting
The quoting rule transfers unchanged — double quotes interpolate, single quotes do not — and format takes the same %-10s and %5d specifications printf does.
name="widget" count=3 echo "double: $count x $name" echo 'single: $name' printf '%-10s|%5d|\n' "$name" "$count"
name = "widget" count = 3 puts "double: #{count} x #{name}" puts 'single: #{name}' puts format('%-10s|%5d|', name, count)
What #{} adds over $name is that the braces are mandatory, so there is no ambiguity about where the name ends and no need for ${name}_suffix. And anything may go inside: #{count * 2}, #{name.upcase}, or a whole method call. Ruby also has %w[] for word arrays and heredocs with <<~TEXT, whose squiggle strips the leading indentation — the thing shell heredocs need <<- and a literal tab for.
Substrings and suffixes
Ruby's path[5, 3] takes an offset and a length, exactly as ${path:5:3} does — one of the few pieces of shell muscle memory that transfers with no adjustment.
path="/var/log/system.log" echo "${path##*/}" # basename echo "${path%/*}" # dirname echo "${path%.log}" # strip a suffix echo "${path:5:3}" # offset and LENGTH
path = "/var/log/system.log" puts File.basename(path) puts File.dirname(path) puts path.delete_suffix(".log") puts path[5, 3] # offset and LENGTH — same as the shell
The shell's expansions are terse and nobody remembers which of # and % works from which end; each one here is a named method instead. Note delete_suffix rather than chomp(".log") — chomp with an argument does the same thing, but strip-shaped methods elsewhere remove any of the given characters repeatedly, which is a real and common bug. For paths, prefer File or Pathname over pattern-matching the string.
Splitting and joining
split and join take the separator as an argument, so nothing global changes. The shell has to reassign IFS, which alters how the whole shell parses words until it is put back.
line="alpha,beta,gamma" IFS=',' read -r -a parts <<< "$line" echo "count: ${#parts[@]}" echo "first: ${parts[0]}" echo "${parts[*]}" | tr ' ' '|'
line = "alpha,beta,gamma" parts = line.split(",") puts "count: #{parts.length}" puts "first: #{parts.first}" puts parts.join("|")
🚨 That IFS dance is the source of a nasty class of bug: forget to scope the assignment and IFS stays changed, so unrelated code five functions away starts splitting on commas. Ruby's separator is also a full regular expression when you want one — line.split(/\s*,\s*/) handles the spaces around the commas in the same call — and parts.first and parts.last retire ${parts[0]} and the ${parts[-1]} that older shells do not support.
Arrays & Hashes
Arrays
🚨 The quotes in "${names[@]}" are not decoration — without them each element is word-split again, so an array holding my report.txt silently becomes two elements when you loop over it. Ruby has nothing to quote.
names=(alpha beta gamma) names+=(delta) echo "count: ${#names[@]}" for name in "${names[@]}"; do echo "- $name" done printf '%s\n' "${names[@]}" | sort -r | tr '\n' ' ' echo
names = %w[alpha beta gamma] names << "delta" puts "count: #{names.length}" names.each { |name| puts "- #{name}" } puts names.sort.reverse.join(" ") + " "
Sorting is the row that shows the size of the change. In the shell, sorting an array means printing it, piping it to sort, and reading the text back — a subprocess and a round trip through a stream. names.sort is a method on the array, it returns a new array, and sort_by { |name| name.length } handles anything sort -k cannot. Ruby arrays also nest, are passed to methods and returned from them, and hold any mixture of types — none of which a Bash array does.
Hashes
declare -A is mandatory and easy to forget: without it Bash evaluates the subscript as arithmetic, an unset word is worth zero, and every key silently lands at index 0.
declare -A counts counts[apple]=2 counts[pear]=5 echo "apple: ${counts[apple]}" for key in "${!counts[@]}"; do echo "$key = ${counts[$key]}" done | sort
counts = { "apple" => 2, "pear" => 5 } puts "apple: #{counts["apple"]}" counts.sort.each do |key, value| puts "#{key} = #{value}" end
Note ${!counts[@]} for the keys against ${counts[@]} for the values — one exclamation mark apart, which is exactly the kind of thing that makes a shell script hard for the next reader. Ruby's block receives both at once. Two shell limits that do not carry over: associative arrays need Bash 4 or later, ruling out macOS's system /bin/bash, and their values can only be strings, so a hash of arrays has no shell equivalent at all.
The structure the shell cannot hold
A hash whose values are hashes — the most ordinary structure there is, and the shell cannot express it. The left column is what people write instead.
# No nesting. The workaround is a flattened key you parse back out, # and it breaks on any value containing the separator you picked. declare -A machine machine["web1:role"]=frontend machine["web1:port"]=8080 machine["db1:role"]=database for key in "${!machine[@]}"; do host="${key%%:*}"; field="${key##*:}" echo "$host $field = ${machine[$key]}" done | sort
machines = { "web1" => { role: "frontend", port: 8080 }, "db1" => { role: "database" }, } machines.sort.each do |host, fields| fields.sort.each do |field, value| puts "#{host} #{field} = #{value}" end end
This is usually the row that ends the argument. Once a script has more than one thing to say about each item, the shell forces you to invent an encoding, and every encoding has a character it cannot survive. Note also that the port on the right is a real integer, comparable and addable, while everything in the left column is text. The role: spelling makes a symbol key — an interned name, cheaper than a string and the conventional choice for a fixed set of field names.
Blocks: The Pipeline, Inside The Language
A pipeline becomes a chain
This is the row the whole page turns on. A chain of methods reads left to right, one stage at a time, exactly as a pipeline does — and tally is the whole of sort | uniq -c.
printf '%s\n' banana apple cherry apple | sort | uniq -c | sort -k1,1nr -k2,2 | awk '{ print $2 ": " $1 }'
fruit = %w[banana apple cherry apple] fruit.tally .sort_by { |name, count| [-count, name] } .each { |name, count| puts "#{name}: #{count}" }
The correspondence is close enough to be worth learning as a table: grep is select, grep -v is reject, awk '{print $2}' is map, sort is sort or sort_by, uniq is uniq, uniq -c is tally, wc -l is count, head is first(n), and xargs is each. What changes is that every stage passes objects rather than text, so nothing is re-parsed between stages and a count is a number you can add to.
What a block actually is
A block is not a loop. It is an argument — a piece of code handed to a method — and the method decides what to do with it, which is why the same syntax covers iterating, transforming, and bracketing a file that must be closed.
# The shell has one thing shaped like this — a loop body — and it # cannot be named, passed anywhere, or reused. for value in 1 2 3; do echo "$(( value * 2 ))" done # Anything else that wants to "do this to each thing" is a pipeline # stage, which means another process and text on both sides.
# A block is a chunk of code passed TO a method. The method decides # when, whether and how often to run it. [1, 2, 3].each { |value| puts value * 2 } doubled = [1, 2, 3].map { |value| value * 2 } puts doubled.inspect # The same block shape brackets a resource, which is what replaces trap: File.open("/tmp/note.txt", "w") { |handle| handle.puts "written" } puts File.read("/tmp/note.txt")
This is the idea with no shell counterpart and the one worth spending an afternoon on, because it explains the rest of Ruby: each, map, select, File.open, Dir.chdir and every timing, retry and transaction wrapper are the same mechanism. The File.open form is the one to notice today — the file is closed when the block ends, including when an exception passes through, which is what trap … RETURN is doing in a shell script.
grep and awk become select and map
Written on three lines so each stage is visible, and it composes the way the pipeline does. The values stay numbers the whole way through instead of being re-parsed from text at every stage.
printf '%s\n' 3 14 15 92 65 | awk '$1 > 10 { print $1 * 2 }'
values = [3, 14, 15, 92, 65] values.select { |value| value > 10 } .map { |value| value * 2 } .each { |value| puts value }
The shorthand you will see everywhere is &: — map(&:upcase) means map { |item| item.upcase }. And when a chain gets long enough to be doing real work, each_with_object, group_by, partition and sum usually replace three stages with one; tally in the first row of this section is the same idea.
Keeping the streaming
.lazy is the word that turns a chain from "build the whole array, then filter it" into "pull one value through at a time" — which is what makes a pipeline able to handle a file bigger than memory.
# The pipeline's real advantage: nothing is ever fully in memory. seq 1 100000 | awk '$1 % 9999 == 0' | head -3
# .lazy makes the chain pull one value at a time, like a pipeline. result = (1..100_000).lazy .select { |value| (value % 9999).zero? } .first(3) puts result
Worth knowing before someone tells you Ruby cannot handle a large file. File.foreach(path) reads one line at a time, each_line and lazy chain the way pipeline stages do, and Enumerator lets you write your own producer. What the shell still wins on is that its stages are separate processes on separate cores, where a lazy chain is one thread doing one thing at a time.
Conditionals & Loops
if, and what counts as true
🚨 Ruby's falsy set is only nil and false. Zero is true. The empty string is true. This is the opposite of both the shell's exit-status convention and Python's falsy-empties rule, and it is the single most important sentence on this page for anyone who already knows another scripting language.
count=0 name="" # Truth is an EXIT STATUS: zero means success. if [[ "$count" -eq 0 ]]; then echo "count is zero"; fi if [[ -z "$name" ]]; then echo "name is empty"; fi if grep -q alpha <<< "alpha beta"; then echo "found it"; fi
count = 0 name = "" text = "alpha beta" # Truth is a property of the VALUE — and only nil and false are false. puts "count is zero" if count.zero? puts "name is empty" if name.empty? puts "found it" if text.include?("alpha")
That rule is a genuine improvement once it clicks: if value means exactly "is there a value here", with no special case for a legitimate zero or an intentionally empty string. The comparison operators simplify too — the shell needs -eq for numbers and == for strings inside [[ ]], where Ruby has one == and "08" == 8 is simply false, because a string is not a number.
case becomes case/in
A comma is the "or", and else is the catch-all that * is. The shapes line up closely enough that a case statement is one of the easier things to port.
for value in start stop restart other; do case "$value" in start) echo "$value: starting" ;; stop|restart) echo "$value: stopping" ;; *) echo "$value: unknown" ;; esac done
%w[start stop restart other].each do |value| case value when "start" then puts "#{value}: starting" when "stop", "restart" then puts "#{value}: stopping" else puts "#{value}: unknown" end end
What when can match is where it stops being a translation. A shell case pattern is a glob against a string; a Ruby when accepts a regular expression, a range (when 1..9), a class (when Integer), or anything defining ===. And case/in, added in Ruby 3.0, destructures instead of comparing: in {role: String => role, port: Integer => port} matches the shape of a hash and binds the pieces in one statement.
Counting loops
(1..3) is a range and includes 3, exactly as {1..3} does — while (1...3), with three dots, stops at 2.
for index in {1..3}; do echo "index $index" done count=0 while (( count < 2 )); do echo "while $count" count=$(( count + 1 )) done
(1..3).each { |index| puts "index #{index}" } count = 0 while count < 2 puts "while #{count}" count += 1 end
One real gain: (1..limit) works with a variable, where the shell's {1..$limit} does not — brace expansion happens before variables are substituted, so it produces the literal text and the loop runs once over nonsense. That is the bug seq exists to work around. Ruby also has 3.times, 1.step(10, 2) and each_slice, and idiomatic code iterates the collection rather than an index.
Methods & Return Values
Parameters get names
The colon after greeting makes it a keyword argument: it is passed by name at the call site, so the reader of greet("world", greeting: "Goodbye") does not have to know what the second position means.
greet() { local name="$1" local greeting="${2:-Hello}" echo "$greeting, $name!" } greet "world" greet "world" "Goodbye"
def greet(name, greeting: "Hello") puts "#{greeting}, #{name}!" end greet("world") greet("world", greeting: "Goodbye")
🚨 local is the most important word in the left column and leaving it out is a silent bug — the variable becomes global and clobbers one three levels up. Ruby has the opposite default: a name assigned inside a method is local to it, full stop, and reaching a script-level variable from inside a method is not even possible without making it global with a $ sigil. Keyword arguments have no shell equivalent at all, and they are what keeps a five-parameter call readable.
🚨 A method returns a value
Ruby returns the last expression evaluated, so return is usually left out entirely. The = form on the first line is an endless method, for a body that is one expression.
# A shell function returns an EXIT STATUS: 0-255, nothing else. double() { echo $(( $1 * 2 )) # the "value" is PRINTED } result="$(double 21)" # and captured by forking a subshell echo "result: $result" # Two values means printing them and splitting them back apart: min_max() { echo "1 9"; } read -r low high <<< "$(min_max)" echo "low=$low high=$high"
def double(value) = value * 2 # the last expression IS the return value result = double(21) puts "result: #{result}" def min_max(values) [values.min, values.max] # an array comes back as an array end low, high = min_max([4, 9, 1]) puts "low=#{low} high=#{high}"
Three costs vanish at once. There is no fork, so a script can be split into small methods without paying for each call. The value keeps its type — a number stays a number, an array stays an array. And returning two things is just returning an array, which the caller destructures with low, high =, rather than an encoding the caller has to reverse. That last one is where shell scripts usually start pushing data through globals and become hard to follow.
All the arguments
"$@" with the quotes is one of the few pieces of shell punctuation worth memorizing: each argument stays one word, spaces and all. $* and unquoted $@ both mangle the second argument here.
show_all() { echo "count: $#" for argument in "$@"; do echo "- [$argument]" done } show_all one "two words" three
def show_all(*arguments) puts "count: #{arguments.length}" arguments.each { |argument| puts "- [#{argument}]" } end show_all("one", "two words", "three")
*arguments collects the rest into an ordinary array, so it can be measured, sliced, passed onward and stored. **options does the same for keyword arguments, which the shell has no concept of, and method(*array) spreads an array back out into separate arguments — the equivalent of writing "${array[@]}" at a call site.
sed, awk and grep
sed s/// becomes sub and gsub
sub replaces the first match and gsub replaces them all — the g that sed puts at the end is the letter at the front. When the pattern is a fixed string, pass a string rather than a regular expression: faster, and it cannot be broken by a special character in the data.
echo "hello world" | sed 's/o/0/g' echo "2026-09-09" | sed -E 's/([0-9]{4})-([0-9]{2})-([0-9]{2})/\3\/\2\/\1/'
puts "hello world".gsub("o", "0") puts "2026-09-09".sub(/(\d{4})-(\d{2})-(\d{2})/, '\3/\2/\1')
Two Ruby-side gains. The replacement can be a block — gsub(/\d+/) { |match| match.to_i * 2 } doubles every number, which sed cannot do at all — and the pattern is not passing through shell quoting first, so there is one layer of escaping rather than two. There is also no delimiter to collide with, so a pattern containing a slash needs no s|…|…| trick.
awk fields become split
Destructuring straight into three named variables is what makes this readable. split with no argument splits on any run of whitespace and drops leading blanks, exactly as awk does by default.
printf '%s\n' "web1 frontend 8080" "db1 database 5432" | awk '{ print $2 " on " $1 ":" $3 }'
["web1 frontend 8080", "db1 database 5432"].each do |line| host, role, port = line.split puts "#{role} on #{host}:#{port}" end
awk is genuinely good at this and a one-line field extraction should stay one line of awk. The case for moving is names — $2 is a mystery on line 40 and role is not — and the fact that the same program can then do something a pipeline cannot: look a value up in a hash, keep a running total in a real number, or write two output files at once. The line where a script stops being a pipeline is usually the line where a stage needs to remember something.
Capturing part of a match
(?<level>…) is a named capture, so the pieces come back by name rather than by number. The match object is returned as a value instead of landing in a global array.
line="ERROR 2026-09-09 disk full" if [[ "$line" =~ ^([A-Z]+)\ ([0-9-]+)\ (.*)$ ]]; then echo "level: ${BASH_REMATCH[1]}" echo "date: ${BASH_REMATCH[2]}" echo "text: ${BASH_REMATCH[3]}" fi
line = "ERROR 2026-09-09 disk full" if (match = line.match(/^(?<level>\w+) (?<date>\S+) (?<text>.*)$/)) puts "level: #{match[:level]}" puts "date: #{match[:date]}" puts "text: #{match[:text]}" end
Three things improve at once: the captures have names, so line 40 of the script is readable; the result is a value, so an unrelated match two lines later cannot clobber it; and the pattern is not passing through the shell's quoting first — notice the backslash-escaped spaces the Bash version needs, and that quoting the pattern inside [[ ]] would silently turn the whole regular expression off.
Files & Paths
Reading and writing a file
Redirection is the shell's best single feature and Ruby has nothing as short. What it has instead is that the file's contents are a value — File.read hands you a string you can do anything with, rather than a stream that has to be consumed once.
printf '%s\n' alpha beta > /tmp/example.txt echo gamma >> /tmp/example.txt cat /tmp/example.txt echo "lines: $(wc -l < /tmp/example.txt | tr -d '[:space:]')"
File.write("/tmp/example.txt", "alpha\nbeta\n") File.open("/tmp/example.txt", "a") { |handle| handle.puts "gamma" } print File.read("/tmp/example.txt") puts "lines: #{File.readlines("/tmp/example.txt").length}"
The block form of File.open closes the handle when the block ends, including when an exception passes through, which is the job trap does in a shell script. For anything small, File.read and File.write skip the block entirely, and File.foreach streams a large file one line at a time. Both columns here write to a filesystem that starts empty on every run, which is why each example creates what it reads.
File tests are the same letters, spelled out
Each cryptic letter becomes a question with its name written on it: -e is exist?, -f is file?, -d is directory?, -s is size?. The question mark is part of the method name and marks it as returning true or false.
printf 'x' > /tmp/probe.txt [[ -e /tmp/probe.txt ]] && echo "exists" [[ -f /tmp/probe.txt ]] && echo "regular file" [[ -d /tmp ]] && echo "directory" [[ -s /tmp/probe.txt ]] && echo "not empty"
File.write("/tmp/probe.txt", "x") puts "exists" if File.exist?("/tmp/probe.txt") puts "regular file" if File.file?("/tmp/probe.txt") puts "directory" if File.directory?("/tmp") puts "not empty" if File.size?("/tmp/probe.txt")
File.size? returns the size rather than merely true, and nil when the file is empty or missing — which is why it works in a condition. 🚨 One warning specific to this page: the permission predicates File.readable?, writable? and executable? answer false in the browser for a file the VM has just written, because the WebAssembly filesystem carries no permission bits for anyone to ask about. Use exist?, file?, directory? and size?, all of which are reliable.
Globbing
Dir.glob takes the same pattern the shell would. The difference that matters is what happens when nothing matches.
mkdir -p /tmp/globdemo : > /tmp/globdemo/one.txt : > /tmp/globdemo/two.txt : > /tmp/globdemo/three.log for file in /tmp/globdemo/*.txt; do echo "found: $(basename "$file")" done
require "fileutils" FileUtils.mkdir_p("/tmp/globdemo") %w[one.txt two.txt three.log].each do |name| File.write("/tmp/globdemo/#{name}", "") end Dir.glob("/tmp/globdemo/*.txt").sort.each do |file| puts "found: #{File.basename(file)}" end
An unmatched shell glob expands to the pattern itself, so the loop body runs once with a filename that does not exist — the bug shopt -s nullglob exists to prevent and that nobody remembers to enable. Dir.glob returns an empty array, so the loop simply does not run. Two more: results come back in filesystem order, so .sort is worth writing when order matters, and Dir.glob("**/*.txt") recurses without needing find, -print0 and xargs -0.
Running Other Programs
Running a command
🚨 The argument list is the whole security story. Open3.capture2("sort", …) hands the arguments straight to the program with no shell in between, so a filename containing a space or a semicolon is a filename. Passing one string instead invokes a shell and reintroduces every injection risk you have ever quoted against.
output="$(printf '%s\n' gamma alpha beta | sort)" echo "$output" echo "status: $?"
require "open3" output, status = Open3.capture2("sort", stdin_data: "gamma\nalpha\nbeta\n") print output puts "status: #{status.exitstatus}" # %x{sort file} and system("sort", "file") also exist. Prefer the # ARRAY form of any of them: no shell means nothing in your data can # be read as a command.
This is where the shell keeps its advantage, and it is worth saying plainly: running programs is what a shell is for, and three lines against one is a real cost. What you get back is that the output, the error output and the status are three separate values you can examine, rather than one string plus a magic $? that the next command overwrites. Open3.capture3 adds standard error as a third value.
Two programs, piped together
A pipe is one character in the shell and a method call in Ruby. This is the clearest case where the shell is simply the better tool, and the honest advice is to notice that rather than argue with it.
printf '%s\n' 3 1 4 1 5 | sort -n | uniq | tr '\n' ' ' echo
require "open3" output, = Open3.pipeline_r(["sort", "-n"], ["uniq"], stdin_data: "3\n1\n4\n1\n5\n") { |out, _| out.read } puts output.split("\n").join(" ") # But this particular pipeline is twenty characters of Ruby: # puts [3, 1, 4, 1, 5].sort.uniq.join(" ")
The comment at the bottom is the real lesson: once the data is inside the program, there is usually no reason to send it back out to a utility at all. sort.uniq does what sort -n | uniq does, without two processes, and it sorts numbers numerically without needing -n. Keep a script in the shell when it is mostly this; move it when the pipeline is a small part of something larger.
Errors & Exit Codes
set -e becomes not needing set -e
An error in Ruby stops the script by default and there is no flag to set. The shell's default is the opposite — a failed command sets a status nobody reads and execution continues — which is what set -e is trying to patch over.
set -euo pipefail # Without those options this script carries on after the failure # and reports success at the end. if ! false; then echo "the command failed, and we noticed" fi echo "still running"
def might_fail raise ArgumentError, "the command failed" end begin might_fail rescue ArgumentError => error puts "#{error.message}, and we noticed" end puts "still running"
set -euo pipefail is the right first line of any shell script and it is still not enough: -e is famously full of exceptions, ignoring failures inside an if condition, in anything followed by &&, and in most of a function called in a condition. Nobody holds the whole rule in their head. Ruby's version is one sentence — an exception propagates until something rescues it — and rescue can be selective by class, which an exit status cannot.
trap becomes ensure — or a block
ensure is trap … RETURN: it runs whether the method returned normally or an exception passed through. Note it is part of the method body — no begin needed.
work() { local scratch scratch="$(mktemp)" trap 'rm -f "$scratch"; echo "cleaned up"' RETURN echo "working" } work
def work path = "/tmp/scratch-#{Process.pid}.txt" File.write(path, "") puts "working" ensure File.delete(path) if path && File.exist?(path) puts "cleaned up" end work # Better still: File.open(path) { … } closes the file at the end of # the block, so most cleanup never needs writing at all.
The comment at the bottom is the idiom to reach for first. Anything that has to be released — a file, a directory you changed into, a lock, a timer — is usually handed out by a method that takes a block and cleans up when the block ends, so the cleanup lives in the library rather than at every call site. That is the same reason the language never needed a finally keyword.
Exiting with a status
Both write the diagnostic to standard error — >&2 on the left, warn on the right — so a caller can separate a message from a result.
check() { if [[ "$1" == "bad" ]]; then echo "fatal: bad input" >&2 return 2 fi echo "ok" } check good check bad || echo "caller saw status $?"
def check(value) raise ArgumentError, "bad input" if value == "bad" puts "ok" end check("good") begin check("bad") rescue ArgumentError => error warn "fatal: #{error.message}" puts "caller saw: #{error.message}" # exit(2) here would hand that status back to the shell. end
An exit status is one number out of 256 and the message is unstructured text on another stream, so a caller can only tell failures apart by convention. An exception carries a class, a message, any attributes you gave it and the stack it came from, so rescue ArgumentError and rescue Errno::ENOENT respond differently without parsing anything. exit(2) still exists for the boundary where a Ruby script genuinely has to answer a shell.
Arguments & One-Liners
Positional arguments
ARGV is an ordinary array holding the arguments — and unlike the shell it does not include the program name, which lives in $PROGRAM_NAME instead.
bash script.sh alpha beta
echo "count: $#" echo "first: $1" echo "all: $*"
ruby main.rb alpha beta
puts "count: #{ARGV.length}" puts "first: #{ARGV.first}" puts "all: #{ARGV.join(" ")}"
Being an array rather than a set of numbered variables is the practical difference: it can be sliced, passed to a method, and iterated without the quoting rules that make "$@" and $* behave differently. For anything with options, do not parse it by hand — the next row is what you want.
getopts becomes OptionParser
Each option is declared once with its short form, its long form and its help text together, and --help is generated from those declarations rather than written separately.
bash script.sh -v -o report.txt
verbose=0 output="" while getopts "vo:" option; do case "$option" in v) verbose=1 ;; o) output="$OPTARG" ;; *) echo "usage: report.sh [-v] [-o FILE]" >&2; exit 1 ;; esac done echo "verbose=$verbose output=${output:-none}" # Short options only. --output means writing the loop yourself.
ruby main.rb --verbose --output report.txt
require "optparse" settings = { verbose: false, output: nil } parser = OptionParser.new do |options| options.banner = "usage: report.rb [options]" options.on("-v", "--verbose", "print more detail") { settings[:verbose] = true } options.on("-o", "--output FILE", "where to write the report") do |file| settings[:output] = file end end parser.parse! # reads the options out of ARGV puts "verbose=#{settings[:verbose] ? 1 : 0} output=#{settings[:output] || "none"}"
This is the row that quietly justifies the move for anything another person will run. getopts handles single letters and nothing else, so every script accepting --dry-run contains a hand-written loop with its own bugs and a usage string that drifts out of date the first time an option is added. OptionParser is in the standard library, needs nothing installed, and also handles type conversion, required arguments and a list of permitted values.
ruby -ne is a pipeline stage
Ruby is a perfectly good pipeline stage, and the flags are the same ones Perl established: -n wraps your program in a loop over input lines, -a splits each into $F, -p also prints. That combination is awk, with Ruby's methods available.
printf '%s\n' "alpha 3" "beta 14" "gamma 15" | awk '$2 > 10 { print $1 }'
# As a one-liner, which is how this gets used: # ... | ruby -ane 'puts $F[0] if $F[1].to_i > 10' # # -e the program -n loop over input lines # -a autosplit into $F -p like -n but prints each line # -i edit files in place # # The same thing as a script: ["alpha 3", "beta 14", "gamma 15"].each do |line| name, value = line.split puts name if value.to_i > 10 end
So moving to Ruby does not mean abandoning the shell — a Ruby one-liner sits exactly where awk did, and reaches for a hash or a standard-library module when the job outgrows one expression. That makes the migration gradual: replace the one stage that has become unreadable, leave the rest, and keep going only as far as it helps. ruby -i.backup -pe is sed -i.
JSON & Structured Data
jq becomes JSON.parse
jq is excellent and is the right answer for a one-off query at a terminal. What it cannot do is hand a value back as anything but text, so everything it produces has to be re-parsed by whatever uses it next.
read -r -d '' payload <<'JSON' || true {"host": "web1", "ports": [80, 443], "tags": {"role": "frontend"}} JSON echo "$payload" | jq -r '.host' echo "$payload" | jq -r '.ports[1]' echo "$payload" | jq -r '.tags.role'
require "json" payload = '{"host": "web1", "ports": [80, 443], "tags": {"role": "frontend"}}' data = JSON.parse(payload) puts data["host"] puts data["ports"][1] puts data["tags"]["role"]
JSON.parse returns real objects: data["ports"] is an array of integers you can sum, sort or append to, and data["ports"][1] + 1 is 444 rather than a string concatenation. Pass symbolize_names: true to get symbol keys, which read better in the rest of the program. The reverse direction is the next row, and it is where the shell struggles most.
Producing JSON safely
Look at the first line of output on the left: the embedded quotation marks break the JSON, silently, and only for data that contains them. This is the most common way a shell script produces output that a downstream consumer rejects.
name='say "hello"' count=3 # Concatenating by hand produces INVALID json the moment a value # contains a quote, a newline or a backslash: echo "{\"name\": \"$name\", \"count\": $count}" # jq -n is the correct way, and it is worth the extra words: jq -nc --arg name "$name" --argjson count "$count" \ '{name: $name, count: $count}'
require "json" name = 'say "hello"' count = 3 puts({ name: name, count: count }.to_json) puts({ nested: { values: [1, 2, 3] } }.to_json)
The shell rule is that JSON must be built with jq -n and --arg for every value, never by interpolation — a rule that gets skipped under deadline pressure precisely because the naive version works on the test data. to_json takes objects rather than text, so quoting is not a decision the caller makes and cannot be got wrong. The same argument applies to building SQL, HTML, URLs and shell commands.
When The Script Grows Up
source becomes require
This is the difference that decides whether a script can grow past one file. source is textual inclusion into one namespace; require_relative loads a file once and a module gives what it defines a name of its own.
# One shared global namespace. Two helper files that both define # "log" silently overwrite each other, and which one wins depends # on the order of the source lines. # source ./helpers.sh log() { echo "[log] $*"; } log "started"
# require_relative "helpers" loads a file once, however many files # ask for it, and a module gives its methods a namespace: # # Helpers.log("started") module Helpers def self.log(*parts) puts "[log] #{parts.join(" ")}" end end Helpers.log("started")
A file is loaded once no matter how many others require it, so there is no include-guard problem, and the require lines at the top are an honest inventory of what a file depends on — which a pile of source lines is not. Two modules may both define log without either being disturbed. A shell script that has reached three files is telling you something.
Testing
Minitest ships with Ruby, so a test file needs nothing installed. Requiring minitest/autorun is what makes the tests run when the file is executed.
# No assertion and no runner in the box. Either a third-party # framework (bats, shunit2) or this: assert_equal() { if [[ "$1" != "$2" ]]; then echo "FAIL: expected [$2], got [$1]" >&2 return 1 fi echo "ok" } double() { echo $(( $1 * 2 )); } assert_equal "$(double 21)" "42"
require "minitest" def double(value) = value * 2 class DoubleTest < Minitest::Test def test_doubles_a_number = assert_equal(42, double(21)) def test_handles_zero = assert_equal(0, double(0)) end Minitest.run
Due to limitations of running Ruby inside the browser, minitest’s autorun feature cannot be used, which merely triggers Minitest.run when the code exits. Instead we must explicitly call Minitest.run. For that reason this code will look slightly different from what you’ll normally see or write. Testability is the quiet reason scripts get rewritten. A shell function can only be tested through the text it prints, so the test asserts on formatting as well as behavior and breaks when a message is reworded. A method returns a value, so the test checks the value and ignores the presentation. The wider tooling assumes that shape too — RSpec, SimpleCov for coverage, RuboCop for style, and Rake for running any of it.
Rake is the shell script you were writing
Every shell script that grows past one job grows this dispatch block. Rake is that block as a library, and the difference is the word deploy: :build — a declared dependency rather than a call.
# The dispatch block every mature shell script grows: run_build() { echo "building"; } run_deploy() { run_build; echo "deploying"; } case "${1:-build}" in build) run_build ;; deploy) run_deploy ;; *) echo "usage: run.sh [build|deploy]" >&2; exit 1 ;; esac
# In a Rakefile, and then: rake deploy # # task :build do # puts "building" # end # # task deploy: :build do # a DEPENDENCY, not a call # puts "deploying" # end # # rake -T lists every task with its description. Dependencies run # once each, in order, however many tasks ask for them. puts "building" puts "deploying"
That distinction earns its keep as soon as three tasks share a setup step: a dependency runs once no matter how many tasks name it, where a called function runs every time. rake -T lists the tasks with their descriptions, which is the usage message you no longer maintain by hand. And a Rakefile is ordinary Ruby, so a task can loop, read a configuration file, or be generated — none of which a case statement can do.
When to stay in the shell
A page like this one can read as an argument that shell scripts are bad. They are not, and the last row should say so plainly: the left column is exactly right as it stands.
#!/usr/bin/env bash # This should NOT be rewritten. It is glue: every line runs a # program, which is exactly what a shell is for. set -euo pipefail mkdir -p /tmp/release printf '%s\n' one two three > /tmp/release/manifest.txt sort /tmp/release/manifest.txt | tee /tmp/release/sorted.txt | wc -l | tr -d '[:space:]'
# The same job in Ruby: longer, no clearer, and it gives up the one # thing the shell is best at. Rewrite when the script has grown LOGIC # — nested data, arithmetic, error handling that must tell failures # apart, several options — not merely because it has grown long. require "fileutils" FileUtils.mkdir_p("/tmp/release") File.write("/tmp/release/manifest.txt", "one\ntwo\nthree\n") lines = File.readlines("/tmp/release/manifest.txt").sort File.write("/tmp/release/sorted.txt", lines.join) puts lines.length
The test is not length. It is whether the script has logic — data with more than one level, arithmetic beyond counting, error handling that must distinguish failures, more than a few options, or anything the next person will modify under pressure. Orchestration and glue belong in the shell. And the move need not be all at once: ruby -ne is a pipeline stage, so the usual path is to replace the one unreadable stage and leave everything around it alone.

Thank you — anything else?