Overview

I used rsync and tar for backups for years. They work, but every time I needed to restore something, I held my breath. Did the incremental sync catch everything? Is this snapshot complete? Did the last tar command actually finish or did I interrupt it?

restic replaces all of that with something that feels closer to Git than to traditional backup tools. Every backup is a snapshot, deduplicated against every other snapshot, encrypted, and independently verifiable. Restores are predictable. Checking integrity is a single command.

Why restic specifically

ToolBackup modelVerificationEncryption
rsyncMirror (not snapshots)None built-inNone
tar + cronFull archivesManual checksumManual (gpg)
borgbackupDeduplicated snapshotsborg checkBuilt-in
resticDeduplicated snapshotsrestic checkBuilt-in

borg and restic are the two serious open-source options. borg has more features and is faster for local backups. restic's advantage is that it treats the repository as dumb storage — S3, Backblaze B2, SFTP, rclone — and doesn't require a server running on the target. You can point it at an S3 bucket and it just works.

That's what decided it for me. Borg wants SSH access to a machine running borg. restic wants a URL.

Installation

# Debian / Ubuntu
sudo apt install restic

# macOS
brew install restic

# Anywhere — static binary
curl -L https://GitHub.com/restic/restic/releases/latest/download/restic_0.17.3_linux_amd64.bz2 | bunzip2 > restic
chmod +x restic
sudo mv restic /usr/local/bin/

Latest releases at restic's GitHub releases page.

Initializing a repository

export RESTIC_REPOSITORY="s3:s3.amazonaws.com/my-backup-bucket"
export RESTIC_PASSWORD="a-very-long-passphrase"

restic init

Two things to internalize now:

  • The password is the only way in. There is no recovery. Lose it, lose your data. Store it in a password manager and somewhere offline.
  • The repository is just a directory structure of blobs. You can look at it with ls and it means nothing. All metadata is encrypted.

For automated backups, use a password file rather than an environment variable:

# /root/.restic-password (mode 600)
cat > /root/.restic-password <<'EOF'
a-very-long-passphrase
EOF

chmod 600 /root/.restic-password
export RESTIC_PASSWORD_FILE=/root/.restic-password

The four commands you'll use

# Back up a directory
restic backup /home/user/documents

# List snapshots
restic snapshots

# Restore a specific snapshot
restic restore latest --target /tmp/restore

# Verify repository integrity
restic check

That's 90% of daily use. Everything else is variations.

Backing up with tags and paths

Tags let you organize snapshots and restore selectively:

restic backup \
  --tag daily \
  --tag home \
  /home/user \
  /etc \
  --exclude-caches \
  --exclude '*.tmp' \
  --exclude '/home/user/.cache'

Then restore only from a specific tag:

restic restore latest --tag daily --target /tmp/restore

Excluding caches matters. --exclude-caches skips directories with a CACHEDIR.TAG file, which many applications create. This alone saves gigabytes on a typical home directory.

How deduplication works in practice

restic chunks files into variable-length blobs and hashes each one. Identical chunks are stored once, regardless of which file or snapshot they came from. On a second backup, only new chunks are uploaded.

The practical consequence: keeping 30 daily snapshots of a 50GB directory uses roughly the space of one full backup plus the daily changes. My typical retention — 7 daily, 4 weekly, 12 monthly, 3 yearly — costs about 1.3x the size of the source data, after a year of accumulation.

Retention policies

restic forget \
  --keep-daily 7 \
  --keep-weekly 4 \
  --keep-monthly 12 \
  --keep-yearly 3 \
  --prune
FlagKeeps
--keep-last NThe N most recent snapshots
--keep-daily NOne snapshot per day for the last N days
--keep-weekly NOne per week for N weeks
--keep-monthly NOne per month for N months
--keep-yearly NOne per year for N years
--keep-within 30dEverything from the last 30 days

--prune is the operation that actually reclaims space. forget alone marks snapshots as removed but doesn't delete the underlying blobs. Running prune after forget in a single command is convenient; separating them makes more sense on large repos where prune takes a long time.

For S3 backends, prune requires reading the entire repository index, which can be slow. Some people run forget frequently (daily) and prune less often (weekly).

Backing up to S3 or B2

export RESTIC_REPOSITORY="s3:https://s3.us-west-004.backblazeb2.com/my-bucket"
export AWS_ACCESS_KEY_ID="..."
export AWS_SECRET_ACCESS_KEY="..."

restic init
restic backup /home/user/documents

Backblaze B2 is the common choice for personal backups — cheap storage, no egress fees at their tier, and it works with the S3 API. If you're using B2's native API (not the S3-compatible one), the repository URL is:

export RESTIC_REPOSITORY="b2:my-bucket:path/to/repo"

I've used both. The S3-compatible endpoint is slightly slower but has fewer moving parts.

Backups on a schedule

A systemd timer is cleaner than cron for this, because it handles failures and logs properly:

# /etc/systemd/system/restic-backup.service
[Unit]
Description=Restic backup
After=network-online.target

[Service]
Type=oneshot
EnvironmentFile=/etc/restic/env
ExecStart=/usr/local/bin/restic backup /home/user /etc
ExecStartPost=/usr/local/bin/restic forget --keep-daily 7 --keep-weekly 4 --keep-monthly 12 --prune
# /etc/systemd/system/restic-backup.timer
[Unit]
Description=Daily restic backup

[Timer]
OnCalendar=daily
Persistent=true
RandomizedDelaySec=1h

[Install]
WantedBy=timers.target
sudo systemctl enable --now restic-backup.timer
sudo systemctl list-timers restic-backup.timer

Persistent=true means a missed backup runs when the machine comes back up. That's important for a laptop that isn't always on.

Verification: the part everyone skips

A backup you haven't verified is a rumor. restic has three levels of checking:

# Quick check — verifies structure and metadata
restic check

# Read all data and verify checksums — slow but thorough
restic check --read-data

# Verify a random 5% of data
restic check --read-data-subset=5%

The default restic check verifies the repository structure but doesn't read all the data. It catches missing blobs and metadata corruption but not silent bit rot in the underlying storage. --read-data-subset=5% on a schedule gives you coverage without the full cost of reading everything.

Run restic check weekly, and --read-data-subset=10% monthly. Then do a real restore once a quarter:

restic restore latest --target /tmp/verify
diff -r /tmp/verify/home/user/documents /home/user/documents

Restoring to a temp directory and diffing against the live data is the only way to know the backup actually contains what you think it does. I've caught a misconfigured exclude pattern this way that would have made every backup useless for six months.

What I'd tell someone starting out

  • Set up the password file before the first backup. Losing the password means losing everything, and it's easy to forget when you're tired at 2am.
  • Start with restic backup only. Add forget and prune once the basic backup works.
  • Restore something on day one. Not a month later. The first restore is the one that catches the config mistakes.
  • Use a repository password manager, and write it on paper. Something offline, somewhere safe. The one time you need it, you won't have access to the manager.
  • Don't back up to the same disk. Obvious, but people do it.

The tool is boring in the way good infrastructure is boring. It runs daily, takes a couple of minutes, and I don't think about it except when I check the timer status. That's the goal.