Overview
I used rsync and tar for backups for years. They work, but every time I needed to restore something, I held my breath. Did the incremental sync catch everything? Is this snapshot complete? Did the last tar command actually finish or did I interrupt it?
restic replaces all of that with something that feels closer to Git than to traditional backup tools. Every backup is a snapshot, deduplicated against every other snapshot, encrypted, and independently verifiable. Restores are predictable. Checking integrity is a single command.
Why restic specifically
| Tool | Backup model | Verification | Encryption |
|---|---|---|---|
| rsync | Mirror (not snapshots) | None built-in | None |
| tar + cron | Full archives | Manual checksum | Manual (gpg) |
| borgbackup | Deduplicated snapshots | borg check | Built-in |
| restic | Deduplicated snapshots | restic check | Built-in |
borg and restic are the two serious open-source options. borg has more features and is faster for local backups. restic's advantage is that it treats the repository as dumb storage — S3, Backblaze B2, SFTP, rclone — and doesn't require a server running on the target. You can point it at an S3 bucket and it just works.
That's what decided it for me. Borg wants SSH access to a machine running borg. restic wants a URL.
Installation
# Debian / Ubuntu
sudo apt install restic
# macOS
brew install restic
# Anywhere — static binary
curl -L https://GitHub.com/restic/restic/releases/latest/download/restic_0.17.3_linux_amd64.bz2 | bunzip2 > restic
chmod +x restic
sudo mv restic /usr/local/bin/
Latest releases at restic's GitHub releases page.
Initializing a repository
export RESTIC_REPOSITORY="s3:s3.amazonaws.com/my-backup-bucket"
export RESTIC_PASSWORD="a-very-long-passphrase"
restic init
Two things to internalize now:
- The password is the only way in. There is no recovery. Lose it, lose your data. Store it in a password manager and somewhere offline.
- The repository is just a directory structure of blobs. You can look at it with
lsand it means nothing. All metadata is encrypted.
For automated backups, use a password file rather than an environment variable:
# /root/.restic-password (mode 600)
cat > /root/.restic-password <<'EOF'
a-very-long-passphrase
EOF
chmod 600 /root/.restic-password
export RESTIC_PASSWORD_FILE=/root/.restic-password
The four commands you'll use
# Back up a directory
restic backup /home/user/documents
# List snapshots
restic snapshots
# Restore a specific snapshot
restic restore latest --target /tmp/restore
# Verify repository integrity
restic check
That's 90% of daily use. Everything else is variations.
Backing up with tags and paths
Tags let you organize snapshots and restore selectively:
restic backup \
--tag daily \
--tag home \
/home/user \
/etc \
--exclude-caches \
--exclude '*.tmp' \
--exclude '/home/user/.cache'
Then restore only from a specific tag:
restic restore latest --tag daily --target /tmp/restore
Excluding caches matters. --exclude-caches skips directories with a CACHEDIR.TAG file, which many applications create. This alone saves gigabytes on a typical home directory.
How deduplication works in practice
restic chunks files into variable-length blobs and hashes each one. Identical chunks are stored once, regardless of which file or snapshot they came from. On a second backup, only new chunks are uploaded.
The practical consequence: keeping 30 daily snapshots of a 50GB directory uses roughly the space of one full backup plus the daily changes. My typical retention — 7 daily, 4 weekly, 12 monthly, 3 yearly — costs about 1.3x the size of the source data, after a year of accumulation.
Retention policies
restic forget \
--keep-daily 7 \
--keep-weekly 4 \
--keep-monthly 12 \
--keep-yearly 3 \
--prune
| Flag | Keeps |
|---|---|
--keep-last N | The N most recent snapshots |
--keep-daily N | One snapshot per day for the last N days |
--keep-weekly N | One per week for N weeks |
--keep-monthly N | One per month for N months |
--keep-yearly N | One per year for N years |
--keep-within 30d | Everything from the last 30 days |
--prune is the operation that actually reclaims space. forget alone marks snapshots as removed but doesn't delete the underlying blobs. Running prune after forget in a single command is convenient; separating them makes more sense on large repos where prune takes a long time.
For S3 backends, prune requires reading the entire repository index, which can be slow. Some people run forget frequently (daily) and prune less often (weekly).
Backing up to S3 or B2
export RESTIC_REPOSITORY="s3:https://s3.us-west-004.backblazeb2.com/my-bucket"
export AWS_ACCESS_KEY_ID="..."
export AWS_SECRET_ACCESS_KEY="..."
restic init
restic backup /home/user/documents
Backblaze B2 is the common choice for personal backups — cheap storage, no egress fees at their tier, and it works with the S3 API. If you're using B2's native API (not the S3-compatible one), the repository URL is:
export RESTIC_REPOSITORY="b2:my-bucket:path/to/repo"
I've used both. The S3-compatible endpoint is slightly slower but has fewer moving parts.
Backups on a schedule
A systemd timer is cleaner than cron for this, because it handles failures and logs properly:
# /etc/systemd/system/restic-backup.service
[Unit]
Description=Restic backup
After=network-online.target
[Service]
Type=oneshot
EnvironmentFile=/etc/restic/env
ExecStart=/usr/local/bin/restic backup /home/user /etc
ExecStartPost=/usr/local/bin/restic forget --keep-daily 7 --keep-weekly 4 --keep-monthly 12 --prune
# /etc/systemd/system/restic-backup.timer
[Unit]
Description=Daily restic backup
[Timer]
OnCalendar=daily
Persistent=true
RandomizedDelaySec=1h
[Install]
WantedBy=timers.target
sudo systemctl enable --now restic-backup.timer
sudo systemctl list-timers restic-backup.timer
Persistent=true means a missed backup runs when the machine comes back up. That's important for a laptop that isn't always on.
Verification: the part everyone skips
A backup you haven't verified is a rumor. restic has three levels of checking:
# Quick check — verifies structure and metadata
restic check
# Read all data and verify checksums — slow but thorough
restic check --read-data
# Verify a random 5% of data
restic check --read-data-subset=5%
The default restic check verifies the repository structure but doesn't read all the data. It catches missing blobs and metadata corruption but not silent bit rot in the underlying storage. --read-data-subset=5% on a schedule gives you coverage without the full cost of reading everything.
Run restic check weekly, and --read-data-subset=10% monthly. Then do a real restore once a quarter:
restic restore latest --target /tmp/verify
diff -r /tmp/verify/home/user/documents /home/user/documents
Restoring to a temp directory and diffing against the live data is the only way to know the backup actually contains what you think it does. I've caught a misconfigured exclude pattern this way that would have made every backup useless for six months.
What I'd tell someone starting out
- Set up the password file before the first backup. Losing the password means losing everything, and it's easy to forget when you're tired at 2am.
- Start with
restic backuponly. Add forget and prune once the basic backup works. - Restore something on day one. Not a month later. The first restore is the one that catches the config mistakes.
- Use a repository password manager, and write it on paper. Something offline, somewhere safe. The one time you need it, you won't have access to the manager.
- Don't back up to the same disk. Obvious, but people do it.
The tool is boring in the way good infrastructure is boring. It runs daily, takes a couple of minutes, and I don't think about it except when I check the timer status. That's the goal.
