Skip to content

Detecting Bash Script Failures with strict mode and trap

Depending on how a Bash script is written, subsequent processing may continue even when a command fails partway through. References to undefined variables and failures in the middle of pipelines can also go unnoticed unless they are handled explicitly.

Common settings that make these failures easier to detect are errexit, nounset, and pipefail. The following combination is commonly used as strict mode.

  • set -e: Exit the shell when an unhandled failure occurs
  • set -u: Treat references to undefined variables as errors
  • set -o pipefail: Reflect failures within a pipeline in its exit status

This article also adds -E, which makes it easier for an ERR trap to be inherited by functions and similar contexts, and uses set -Eeuo pipefail as the basic form.

However, set -e does not mean that the shell will always stop on every nonzero exit status. Its behavior differs when failures are evaluated by control structures, such as conditions in if, while, or until, parts of && or || lists, and commands whose status is inverted with !. Therefore, you should not rely on strict mode alone; expected failures must be handled explicitly by design.

Step 1: Stop unhandled failures with errexit

Section titled “Step 1: Stop unhandled failures with errexit”

When set -e is enabled, if a command returns a nonzero status on the normal execution path, the shell can terminate instead of continuing to subsequent processing.

In the following example, false returns exit status 1. after is not executed, and the inner Bash process exits with a nonzero status.

Terminal window
status=0
output=$(bash -c '
set -e
printf "%s\n" "before"
false
printf "%s\n" "after"
' 2>&1) || status=$?
printf '%s\n' "$output"
printf 'exit=%d\n' "$status"
test "$status" -ne 0
! grep -qx 'after' <<<"$output"

This prevents a script from ignoring an error and continuing into subsequent update or deletion operations.

When a script intentionally inspects exit statuses, however, you must also understand the exception conditions of set -e.

Step 2: Detect undefined variables with nounset

Section titled “Step 2: Detect undefined variables with nounset”

A typo in a shell variable name or a missed initialization can be difficult to diagnose if processing continues with an empty string.

When set -u is enabled, references to undefined variables are treated as errors.

Terminal window
status=0
output=$(bash -c '
unset BASH_STRICT_DEMO_UNSET
set -u
printf "%s\n" "before"
printf "%s\n" "$BASH_STRICT_DEMO_UNSET"
printf "%s\n" "after"
' 2>&1) || status=$?
printf '%s\n' "$output"
printf 'exit=%d\n' "$status"
test "$status" -ne 0
! grep -qx 'after' <<<"$output"

As soon as the undefined variable is referenced, the inner Bash process exits with a nonzero status and does not reach after.

However, directly referencing values that are allowed to be unset can also cause valid cases to fail. For optional arguments and similar values, use a default such as ${1:-default} and distinguish between required values and values that may be unset.

Step 3: Detect failures inside pipelines with pipefail

Section titled “Step 3: Detect failures inside pipelines with pipefail”

For a normal pipeline, the exit status of the last command is generally used as the result of the entire pipeline.

Therefore, in a sequence like the following, even if the false command on the left fails, the entire pipeline is treated as successful when the final true command succeeds.

When pipefail is enabled, if any command in the pipeline fails, that failure can be reflected in the exit status of the entire pipeline.

Terminal window
bash -c '
set +o pipefail
false | true
without_pipefail=$?
set -o pipefail
false | true
with_pipefail=$?
printf "without_pipefail=%d\n" "$without_pipefail"
printf "with_pipefail=%d\n" "$with_pipefail"
test "$without_pipefail" -eq 0
test "$with_pipefail" -ne 0
'

This is particularly important in scripts that connect multiple commands with pipes, such as sending logs to grep or compressing and transforming generated output.

Combining set -e with pipefail also makes a failure in the middle of a pipeline eligible to stop the script as an unhandled error.

Step 4: Run failure handling with an ERR trap

Section titled “Step 4: Run failure handling with an ERR trap”

When an ERR trap is configured, diagnostic processing can run when a target command exits with a nonzero status.

Because $? contains the exit status from immediately before the trap was invoked, save it first and then perform logging or other processing.

In the following example, false runs inside a function. Enabling -E causes the ERR trap to be inherited inside the function as well.

cat > /tmp/bash-err-trap-demo.sh <<'BASH'
#!/usr/bin/env bash
set -Ee
trap 'rc=$?; printf "ERR rc=%d\n" "$rc"' ERR
fail_in_function() {
false
printf '%s\n' 'after-function'
}
fail_in_function
printf '%s\n' 'after-script'
BASH
status=0
output=$(bash /tmp/bash-err-trap-demo.sh 2>&1) || status=$?
rm -f /tmp/bash-err-trap-demo.sh
printf '%s\n' "$output"
printf 'exit=%d\n' "$status"
test "$status" -ne 0
grep -qx 'ERR rc=1' <<<"$output"
! grep -qx 'after-function' <<<"$output"
! grep -qx 'after-script' <<<"$output"

An ERR trap is not a general exception-handling mechanism that catches every failure. As with set -e, there are cases where it does not run when the failure is part of a control structure, such as a conditional test.

For that reason, use an ERR trap mainly for diagnosing or logging unexpected failures, and handle expected failures explicitly with conditional branches.

Step 5: Handle expected failures explicitly

Section titled “Step 5: Handle expected failures explicitly”

Even with strict mode enabled, a nonzero exit can be a normal branching condition, such as when another action should be performed if something is not found.

Treat such commands as conditions in if or similar constructs and handle the failure explicitly.

cat > /tmp/bash-expected-failure-demo.sh <<'BASH'
#!/usr/bin/env bash
set -Eeuo pipefail
err_seen=0
trap 'err_seen=1' ERR
if false; then
printf '%s\n' 'unexpected-success'
else
printf '%s\n' 'expected-failure-handled'
fi
printf '%s\n' 'continued'
test "$err_seen" -eq 0
BASH
bash /tmp/bash-expected-failure-demo.sh
rm -f /tmp/bash-expected-failure-demo.sh

In this example, even though false fails, the script does not exit because the command itself is the condition tested by if; execution proceeds to else. The ERR trap is not executed either.

In real scripts, distinguishing the following two categories makes error handling clearer.

  • Unexpected failure: Stop through strict mode and, when necessary, diagnose it with an ERR trap
  • Expected failure: Handle it explicitly with if, ||, or a similar construct

Unconditionally adding || true to the end of a command risks treating failures that should stop the script as successes. If you intentionally ignore a failure, limit both the reason and the scope for doing so.

Step 6: Clean up with an EXIT trap and preserve the exit status

Section titled “Step 6: Clean up with an EXIT trap and preserve the exit status”

Scripts that use temporary directories or files should perform cleanup even if they fail partway through.

An EXIT trap can be used for common processing when the shell exits. It is especially important to save $? when the trap starts and return the original exit status after cleanup.

The following example runs both a successful and a failing case and verifies that the temporary directories are removed while the original exit results are preserved.

cat > /tmp/bash-cleanup-demo.sh <<'BASH'
#!/usr/bin/env bash
set -Eeuo pipefail
workdir=$(mktemp -d /tmp/bash-cleanup-demo.XXXXXX)
cleanup() {
rc=$?
trap - EXIT
rm -rf "$workdir" || true
printf 'cleanup rc=%d\n' "$rc"
exit "$rc"
}
trap cleanup EXIT
printf 'workdir=%s\n' "$workdir"
if [[ ${1:-ok} == fail ]]; then
false
fi
printf '%s\n' 'completed'
BASH
success_output=$(bash /tmp/bash-cleanup-demo.sh)
success_status=$?
success_dir=$(awk -F= '/^workdir=/{print $2}' <<<"$success_output")
failure_status=0
failure_output=$(bash /tmp/bash-cleanup-demo.sh fail 2>&1) || failure_status=$?
failure_dir=$(awk -F= '/^workdir=/{print $2}' <<<"$failure_output")
rm -f /tmp/bash-cleanup-demo.sh
printf '%s\n' "$success_output"
printf 'success_exit=%d\n' "$success_status"
printf '%s\n' "$failure_output"
printf 'failure_exit=%d\n' "$failure_status"
test "$success_status" -eq 0
test "$failure_status" -ne 0
test ! -e "$success_dir"
test ! -e "$failure_dir"
grep -qx 'cleanup rc=0' <<<"$success_output"
grep -qx 'cleanup rc=1' <<<"$failure_output"

At the start of cleanup, the exit status is saved in rc.

By removing the current EXIT trap with trap - EXIT and then calling exit "$rc", the original exit status is returned to the caller after cleanup completes. The rm -rf command also uses || true so that a cleanup failure does not overwrite the original exit status.

The successful path preserves status 0, while the failing path stopped by false preserves a nonzero status.

Writing set -Eeuo pipefail at the beginning does not complete all error handling by itself.

Keeping the following points in mind makes it easier to build safer scripts.

  • Handle cases where a nonzero exit is normal explicitly with if or a similar construct
  • Enable pipefail for pipelines
  • Use forms such as ${VAR:-default} for variables that may be unset
  • Save $? first in an ERR trap
  • Use an EXIT trap to remove temporary resources
  • Preserve the original exit status in the EXIT trap
  • Do not hide errors with an unjustified || true

Also, an EXIT trap performs cleanup only when the shell exits in a way that allows the trap to run. It cannot guarantee cleanup for termination that the process cannot handle, such as SIGKILL.

It is easier to design strict mode correctly when you think of it not as a feature that automatically fixes failures, but as a mechanism that reduces the range of cases in which failures can be missed while processing continues.

Category: Linux