Overview
I built a home NAS on ext4 with mdadm RAID five years ago, and it worked right up until a disk failed, the rebuild hit an unreadable sector on another disk, and I lost 4TB of photos. That was the week I learned why people are so loud about ZFS.
ZFS isn't a filesystem in the usual sense. It's a volume manager, filesystem, RAID implementation, and snapshot engine stacked together, and it treats data integrity as the primary goal rather than an afterthought. That design choice is what makes it worth the learning curve.
What's actually different
| Concern | mdadm + ext4 | ZFS |
|---|---|---|
| Silent bit rot | Undetected | Detected and repaired |
| RAID rebuild | Fails on unreadable sector | Reconstructs from parity |
| Snapshots | None (or LVM, awkward) | Instant, cheap, first-class |
| Compression | No | Transparent, per-dataset |
| RAID level changes | Requires rebuild | Add or replace vdevs |
| Memory usage | Minimal | 1GB per TB of storage + ARC |
| Expansion | Add disks, resize | Add vdevs to a pool |
The silent bit rot one is the reason I switched. ZFS checksums every block and stores a copy of the checksum with the parent block. When you read data, it verifies the checksum. If it doesn't match, ZFS pulls the correct copy from a mirror or parity. You find out because ZFS logs a checksum error, not because your photo opens as a garbled mess three years later.
Installing
Debian and Ubuntu ship ZFS in the kernel module package:
sudo apt install zfsutils-linux
sudo modprobe zfs
zfs --version
On RHEL and derivatives, the OpenZFS repo has current packages. On Arch, zfs-dkms from the AUR. The OpenZFS documentation lists every distro's install path.
One thing to know before installing: ZFS is licensed under the CDDL, which is incompatible with the GPL. It ships as a kernel module rather than being in the kernel tree. This has been fine for a decade, and it means you'll occasionally need to rebuild the module after a kernel upgrade on a distro that doesn't ship prebuilt modules.
Creating a pool
# Four-disk RAID-Z2 (like RAID 6, can lose two disks)
sudo zpool create tank raidz2 /dev/sda /dev/sdb /dev/sdc /dev/sdd
# Two-disk mirror (like RAID 1)
sudo zpool create tank mirror /dev/sda /dev/sdb
# Pool spanning multiple vdevs — this is how you scale
sudo zpool create tank mirror /dev/sda /dev/sdb mirror /dev/sdc /dev/sdd
| Layout | Disks | Fault tolerance | Usable |
|---|---|---|---|
| mirror | 2 | 1 disk | 50% |
| raidz1 | 3+ | 1 disk | (n-1)/n |
| raidz2 | 4+ | 2 disks | (n-2)/n |
| raidz3 | 5+ | 3 disks | (n-3)/n |
Use whole disks (/dev/sda), not partitions. ZFS wants to manage the layout itself, and using partitions means you're fighting it later.
The critical rule: you can add vdevs to a pool, but you cannot add disks to a vdev. A four-disk raidz2 vdev stays a four-disk raidz2 vdev forever. To expand, add another four-disk raidz2 vdev to the pool. Plan your vdev size at creation.
Datasets: the layers above the pool
ZFS pools are just storage. Datasets are what you actually use:
sudo zfs create tank/photos
sudo zfs create tank/documents
sudo zfs create tank/backups
zfs list
# NAME USED AVAIL REFER MOUNTPOINT
# tank 200G 3.4T 96K /tank
# tank/photos 180G 3.4T 180G /tank/photos
# tank/documents 18G 3.4T 18G /tank/documents
# tank/backups 2G 3.4T 2G /tank/backups
Each dataset is a separate filesystem with its own settings and its own snapshots. This is where ZFS shines — you can enable compression on one dataset, disable it on another, snapshot one without affecting the others.
# Enable compression on a dataset (lz4 is the default and best choice)
sudo zfs set compression=lz4 tank/documents
# Set a quota
sudo zfs set quota=100G tank/backups
# Reserve space so other datasets can't consume it
sudo zfs set reservation=500G tank/photos
# Enable deduplication — read the warning below before doing this
sudo zfs set dedup=on tank/vms
Compression is nearly free and you should enable it everywhere. lz4 is fast enough that it rarely slows down I/O, and text, logs, and VMs compress by 2-3x. Databases compress less; media files barely at all.
Deduplication is the trap. It requires a dedup table in RAM — roughly 1-5GB per TB of unique data — and if the table doesn't fit, performance collapses. I've seen people turn it on and then wonder why their NAS reads at 20MB/s. Enable it only if you have a specific use case (backup targets with many similar VMs) and enough RAM to hold the table.
Snapshots: the feature that changes how you work
# Snapshot a dataset
sudo zfs snapshot tank/photos@before-migration
# List snapshots
zfs list -t snapshot
# Roll back to a snapshot
sudo zfs rollback tank/photos@before-migration
# Access snapshot files without rolling back
ls /tank/photos/.zfs/snapshot/before-migration/
# Destroy a snapshot
sudo zfs destroy tank/photos@before-migration
Snapshots are copy-on-write, which means they take almost no space until the data they point to changes. A snapshot of a 500GB dataset with no changes uses a few KB of metadata. This makes it practical to snapshot hourly.
The .zfs/snapshot/ hidden directory is what makes this usable day to day. You don't have to roll back to see old files — you can browse any snapshot as a regular directory. I've recovered dozens of accidentally deleted files this way, without any special tooling.
Automated snapshots with sanoid
sudo apt install sanoid
# /etc/sanoid/sanoid.conf
[tank/photos]
use_template = production
recursive = yes
[tank/documents]
use_template = production
recursive = yes
[template_production]
frequently = 0
hourly = 24
daily = 30
monthly = 3
yearly = 0
autosnap = yes
autoprune = yes
sudo systemctl enable --now sanoid.timer
Sanoid takes snapshots and prunes old ones on a schedule. The autoprune setting is what keeps it from filling the disk — it removes snapshots older than the retention window automatically.
Scrubs: verifying your data
# Start a scrub manually
sudo zpool scrub tank
# Check status
sudo zpool status tank
# Schedule monthly scrubs via cron
0 3 1 * * /sbin/zpool scrub tank
A scrub reads every block, verifies the checksum, and repairs any mismatches using redundancy. This is the operation that actually finds bit rot, and you should run it monthly.
The output tells you how many errors were found and repaired. If zpool status ever shows Permanent errors in the non-zero state, that's data that couldn't be repaired — meaning you lost redundancy for that block. Investigate immediately.
Replacing a failed disk
# Identify the failed disk
sudo zpool status tank
# Physically replace it, then:
sudo zpool replace tank /dev/sda /dev/sde
# Watch progress
watch -n 5 'sudo zpool status tank'
ZFS rebuilds from redundancy. Unlike mdadm, if it hits an unreadable sector on another disk during the rebuild, it can reconstruct it from parity and log a checksum error rather than failing the rebuild. This is the difference that mattered to me when I lost my photos.
Resilver time depends on disk size and pool layout. A 4TB disk in a four-disk raidz2 takes about four to six hours on decent hardware. During that window, performance is degraded but the pool stays online.
Memory and tuning
ZFS uses RAM as a read cache (the ARC) and, if you enable it, as a write cache (the L2ARC, usually an SSD). The rule of thumb is 1GB of RAM per TB of storage, plus whatever the OS needs.
The setting that most people change:
# Limit ARC to 8GB
echo "options zfs zfs_arc_max=8589934592" | sudo tee /etc/modprobe.d/zfs.conf
sudo update-initramfs -u
sudo reboot
By default, ZFS uses up to half of system RAM for the ARC. On a dedicated NAS, that's fine. On a machine running other things, limit it, or it will crowd out your application memory.
When ZFS is the wrong choice
- Single-disk laptops. ZFS on one disk doesn't provide redundancy, and the overhead isn't worth it. Use ext4 or btrfs.
- Machines with less than 8GB of RAM. The ARC needs room, and a constrained ZFS system performs worse than a well-tuned ext4.
- Hardware RAID controllers. ZFS wants raw disk access. Putting it on top of a RAID controller defeats the purpose and can cause corruption if the controller caches writes. Use a plain HBA.
- Windows. There's no production-quality ZFS for Windows. Use Storage Spaces or a NAS.
For a home NAS, a server with multiple disks, or anywhere you care about data integrity, ZFS is the right default. The learning curve is real, but the failure modes of the alternatives are worse.
One last thing: back up your ZFS pool. Redundancy is not backup. Snapshot replication to another machine with zfs send | zfs receive is the correct way, and it's a topic on its own.
