ZFS Deep Dive | TrueNAS - Wyatt's Notes
ZFS Architecture
Section titled “ZFS Architecture”The Three Layers
Section titled “The Three Layers”ZFS is not a traditional filesystem. It is a combined volume manager and filesystem built from three Distinct layers:
graph TD A[Application / POSIX Interface] --> B[ZPL - ZFS POSIX Layer] B --> C[DMU - Data Management Unit] C --> D[SPA - Storage Pool Allocator] D --> E[Physical Storage - vdevs]- ZPL (ZFS POSIX Layer): The filesystem layer that provides POSIX-compliant semantics. Files, directories, permissions, extended attributes, and ACLs. It translates file operations into block operations.
- DMU (Data Management Unit): The transactional layer that manages objects, blocks, and snapshots. All writes are handled as atomic transactions. The DMU also manages the ARC (Adaptive Replacement Cache).
- SPA (Storage Pool Allocator): The lowest layer that manages physical storage. It handles vdev topology, I/O scheduling, checksumming, compression, and self-healing.
Copy-on-Write Transaction Model
Section titled “Copy-on-Write Transaction Model”ZFS never overwrites data in place. Every write creates a new copy of the data block, and only after The new block is written and its checksum verified does ZFS update the metadata to point to the new Block. This has several consequences:
- Snapshots are instantaneous and free (initially). A snapshot is a marker in the transaction history that prevents old blocks from being freed.
- No write hole. Unlike hardware RAID, a power loss during a write cannot leave data and parity in an inconsistent state. Either the old data or the new data is referenced, never a partial update.
- Fragmentation is inevitable. Over time, as blocks are updated and freed, the pool becomes fragmented. This is the primary trade-off of copy-on-write.
Merkle Tree and Checksumming
Section titled “Merkle Tree and Checksumming”Every block in ZFS is identified by its content hash (SHA-256 by default), forming a Merkle tree. The root of the Merkle tree (the block pointer) stores the checksum of its child blocks, and so on Recursively.
When ZFS reads a block, it:
- Reads the block and its checksum from disk.
- Recomputes the checksum of the read data.
- Compares the computed checksum with the stored checksum.
- If they match, the data is returned.
- If they do not match (silent corruption), ZFS uses redundancy (mirror or parity) to reconstruct the correct data and repair the corrupted copy.
Checksum Algorithms
Section titled “Checksum Algorithms”| Algorithm | Speed | Collision Resistance | Recommendation |
|---|---|---|---|
| fletcher2 | Fast | Low | Legacy only |
| fletcher4 | Fast | Low | Default on older pools |
| sha256 | Moderate | High | Default and recommended |
| sha512 | Slow | Very High | Security-sensitive environments |
| edonr | Very Fast | Very High | Best for modern hardware (SSE4.2+) |
| blake3 | Very Fast | Very High | Available on newer ZFS versions |
Storage Pools
Section titled “Storage Pools”vdev Types
Section titled “vdev Types”A ZFS pool (zpool) is constructed from one or more vdevs (virtual devices). The vdev is the Fundamental unit of redundancy and performance. Data is striped across vdevs for performance.
graph TD A[zpool] --> B[vdev 0: mirror-0] A --> C[vdev 1: raidz2-0] B --> D[disk1] B --> E[disk2] C --> F[disk3] C --> G[disk4] C --> H[disk5] C --> I[disk6]| vdev Type | Min Drives | Fault Tolerance | Capacity Efficiency | Write Performance | Read Performance |
|---|---|---|---|---|---|
| stripe (none) | 1 | None | 100% | N × | N × |
| mirror | 2 | N-1 drives | 1/N × | 1 × (per mirror) | N × |
| raidz1 | 3 | 1 drive | (N-1)/N × | Moderate | Good |
| raidz2 | 4 | 2 drives | (N-2)/N × | Moderate | Good |
| raidz3 | 5 | 3 drives | (N-3)/N × | Moderate | Good |
| draid1 | 3 | 1 drive + spare | Similar to raidz1 | Good | Good |
| draid2 | 4 | 2 drives + spare | Similar to raidz2 | Good | Good |
RAIDZ Dynamic Striping
Section titled “RAIDZ Dynamic Striping”RAIDZ uses dynamic stripe width. Unlike traditional RAID 5/6 where the stripe width is fixed (e.g., 4+1 for RAID 5 with 5 drives), RAIDZ varies the stripe width based on the size of the incoming Write:
- Small writes (less than one sector per data disk) are written as “full stripe” writes with variable-width padding.
- Large writes that fill the stripe exactly avoid any padding overhead.
- Medium writes may leave some sectors unused (wasted space).
This is why RAIDZ capacity is slightly less than the theoretical (N-P)/N formula, where P is the Parity count. The actual usable capacity depends on the recordsize and write patterns.
ashift (Sector Size)
Section titled “ashift (Sector Size)”The ashift property controls the physical sector size that ZFS assumes for the drives. It must be Set at pool creation time and cannot be changed afterward.
| ashift | Sector Size | When to Use |
|---|---|---|
| 9 | 512 bytes | Legacy drives only |
| 12 | 4 KB | Most modern HDDs and SSDs |
| 13 | 8 KB | Some modern SSDs with 8 KB physical sectors |
| 14 | 16 KB | Advanced-format SMR drives |
Pool Creation Examples
Section titled “Pool Creation Examples”# Mirror pool (2-way mirror)zpool create tank mirror /dev/sda /dev/sdb
# RAIDZ2 pool with ashift=12zpool create -o ashift=12 tank raidz2 /dev/sda /dev/sdb /dev/sdc /dev/sdd
# Hybrid pool: SSD mirror for special vdev + HDD RAIDZ2 for datazpool create -o ashift=12 tank \ mirror /dev/nvme0n1 /dev/nvme1n1 \ raidz2 /dev/sda /dev/sdb /dev/sdc /dev/sdd /dev/sde /dev/sdfDataset Hierarchy
Section titled “Dataset Hierarchy”Datasets and Properties
Section titled “Datasets and Properties”A dataset (zfs filesystem) is a logical namespace within a pool. Each dataset has its own Properties, mount point, and can have its own snapshots, quotas, and compression settings.
## Create a datasetzfs create tank/data
## Set propertieszfs set compression=lz4 tank/datazfs set atime=off tank/datazfs set recordsize=128K tank/datazfs set quota=500G tank/data
# List all propertieszfs get all tank/dataKey Dataset Properties
Section titled “Key Dataset Properties”| Property | Default | Description | Recommendation |
|---|---|---|---|
| compression | off (on), lz4 (TrueNAS) | Compress data before writing | Always lz4 (fast, low CPU) or zstd (better ratio) |
| atime | on | Update file access time on read | Set off to reduce metadata writes |
| recordsize | 128K | Maximum block size for a file | 128K for media, 16K-64K for VMs, 8K for databases |
| dedup | off | Deduplicate blocks | Generally off (high memory cost) |
| sync | standard | Synchronous write behavior | standard for NFS, disabled for scratch |
| logbias | latency | Optimize for latency vs throughput | latency for databases, throughput for media |
| primarycache | all | What to store in ARC | all for most workloads |
| secondarycache | all | What to store in L2ARC | all if L2ARC present |
recordsize Selection
Section titled “recordsize Selection”The recordsize property determines the maximum block size ZFS uses for a file. ZFS uses Variable-size blocks up to this maximum. The optimal recordsize depends on the workload:
| Workload | Recommended recordsize | Rationale |
|---|---|---|
| Media files (video, audio, images) | 128K (default) | Large sequential reads benefit from large blocks |
| Virtual machine images | 64K or 16K | VMs do mixed random/sequential I/O |
| Databases (MySQL, PostgreSQL) | 8K or 16K | Match the database page size |
| General file storage | 128K | Good balance for mixed workloads |
| NFS home directories | 128K | Mixed workload, default is fine |
Snapshots
Section titled “Snapshots”How Snapshots Work
Section titled “How Snapshots Work”Because ZFS is copy-on-write, a snapshot is a point-in-time marker in the transaction History. Creating a snapshot is instantaneous and consumes no space initially. Space is consumed Only when blocks referenced by the snapshot are modified or deleted in the live filesystem.
Snapshot Space Accounting
Section titled “Snapshot Space Accounting”The space used by a snapshot is the total size of blocks that have been modified or deleted in the Live filesystem since the snapshot was taken. This is called “written” space:
# List snapshots and their space usagezfs list -t snapshot -o name,used,refer,written
# Check space used by a specific snapshotzfs list -o name,used,refer tank/data@daily.2024-01-01Snapshot Lifecycle Management
Section titled “Snapshot Lifecycle Management”# Create a snapshotzfs snapshot tank/data@daily.2024-01-01
# List snapshotszfs list -t snapshot
# Destroy a snapshotzfs destroy tank/data@daily.2024-01-01
# Destroy snapshots matching a patternzfs destroy tank/data@daily.2023-*
# Clone a snapshot (creates a writable copy)zfs clone tank/data@daily.2024-01-01 tank/data-restore
# Promote a clone (make it independent of the snapshot)zfs promote tank/data-restoreClones vs. Snapshots
Section titled “Clones vs. Snapshots”| Feature | Snapshot | Clone |
|---|---|---|
| Writable | No | Yes |
| Space usage | Only changed blocks | Same as snapshot + new writes |
| Can be mounted | No | Yes |
| Dependencies | Cannot destroy if clone exists | Independent after promotion |
| Use case | Backup points, rollback | Testing, temporary environments |
ARC, L2ARC, and SLOG
Section titled “ARC, L2ARC, and SLOG”ARC (Adaptive Replacement Cache)
Section titled “ARC (Adaptive Replacement Cache)”The ARC is ZFS”s primary read cache, stored in system RAM. It uses the Adaptive Replacement Cache Algorithm, which maintains two lists:
- MRU (Most Recently Used): Recently accessed data.
- MFU (Most Frequently Used): Frequently accessed data.
The ARC dynamically balances between these two lists, evicting from the list with lower hit rates. This performs better than a simple LRU cache for mixed workloads with both sequential and random Access patterns.
ARC sizing: The default ARC maximum is 50% of system RAM on TrueNAS. For dedicated NAS Workloads, increasing the ARC to 70–80% of RAM can significantly improve read performance for hot Datasets.
# Check ARC statszfs get arcstats 2>/dev/null || cat /proc/spl/kstat/zfs/arcstats
# Key metrics:# arc_hits — Cache hits# arc_misses — Cache misses# arc_hit_ratio — Percentage of reads served from cacheL2ARC (Level 2 ARC)
Section titled “L2ARC (Level 2 ARC)”The L2ARC is a secondary read cache stored on a dedicated SSD (or partition of an SSD). When the ARC Evicts data, it can write it to the L2ARC before discarding it entirely. On subsequent accesses, if The data is not in the ARC but is in the L2ARC, it can be read from the L2ARC rather than from the Slower pool disks.
L2ARC considerations:
- L2ARC is read-through, not write-through. Data is written to L2ARC only when evicted from ARC.
- The L2ARC does not speed up writes — only reads.
- L2ARC requires significant ARC space to be effective. The ARC metadata for tracking L2ARC entries consumes RAM.
- L2ARC is most effective when the working set is larger than ARC but smaller than ARC + L2ARC.
SLOG (ZIL)
Section titled “SLOG (ZIL)”The SLOG (Separate Log) is an accelerator for synchronous writes. When a synchronous write request Arrives (from NFS, SMB sync, or a database), ZFS must ensure the data is on stable storage before Acknowledging the write. Without a SLOG, this means writing directly to the pool, which is slow for HDD-based pools.
A dedicated SLOG device ( a low-latency NVMe SSD or Intel Optane) absorbs synchronous Writes at SSD speed, then asynchronously flushes them to the pool. This dramatically improves NFS And database write performance on HDD-based pools.
Scrub and Resilver
Section titled “Scrub and Resilver”A scrub reads all data in the pool and verifies checksums. If a checksum mismatch is detected (the Block is corrupted), ZFS automatically repairs it from a redundant copy (mirror or parity).
# Start a scrubzpool scrub tank
# Check scrub statuszpool status tank
# Scrub scheduling on TrueNAS:# Configure under Data Protection → Scrub Tasks# Recommended: Monthly scrubs for HDD pools, Weekly for SSD poolsScrub best practices:
- Run scrubs at off-peak hours. Scrubbing a large HDD pool can take days and significantly impacts pool performance.
- SSD pools scrub much faster (hours instead of days) due to higher throughput.
- Monitor scrub progress with
zpool status. The scrub will report any errors found and repaired. - If a scrub finds uncorrectable errors, immediately back up critical data and replace the failing drive.
Resilver
Section titled “Resilver”A resilver rebuilds the data on a replaced drive. Unlike traditional RAID rebuilds, ZFS resilvers Only copy the actual data (not the entire disk), and they prioritize data based on its metadata Importance.
# Replace a failed drivezpool replace tank /dev/sda /dev/sdb
# Monitor resilver progresszpool status tankSend and Receive
Section titled “Send and Receive”Incremental Replication
Section titled “Incremental Replication”ZFS send/receive is the native mechanism for replicating datasets between pools or systems. It works At the block level, sending only the changed blocks between two snapshots.
# Full replication (initial)zfs send tank/data@snapshot1 | zfs recv backup/data
# Incremental replication (send only changes since snapshot1)zfs send -i tank/data@snapshot1 tank/data@snapshot2 | zfs recv backup/data
# Replication with compression over SSHzfs send -Rcv tank/data@snapshot1 | ssh nas2 zfs recv backup/data
# Raw send (preserves encryption and compression)zfs send -w tank/data@snapshot1 | zfs recv backup/dataReplication Strategies
Section titled “Replication Strategies”| Strategy | Bandwidth | Storage | Complexity |
|---|---|---|---|
| Full periodic | High | High | Low |
| Incremental periodic | Low | Medium | Medium |
| Continuous (zfs-auto-snapshot + cron) | Low | Medium | Medium |
| TrueNAS replication task | Low | Medium | Low (GUI) |
zpool Status Interpretation
Section titled “zpool Status Interpretation”Reading zpool status
Section titled “Reading zpool status”zpool status -v tankKey fields to understand:
| Field | Meaning |
|---|---|
| state | Overall pool state (ONLINE, DEGRADED, FAULTED, UNAVAIL) |
| status | Human-readable description of current state |
| action | Recommended corrective action |
| see | Kernel message log reference |
| config | Detailed vdev and disk status |
| errors | Read, write, and checksum error counts per disk |
Drive Status Values
Section titled “Drive Status Values”| Status | Meaning | Action |
|---|---|---|
| ONLINE | Drive is healthy and active | None |
| DEGRADED | Drive is operational but pool redundancy is reduced | Replace failed drive |
| OFFLINE | Drive has been taken offline administratively | Bring online or replace |
| FAULTED | Drive has been marked as failed | Replace immediately |
| UNAVAIL | Drive cannot be opened or accessed | Check connections, replace |
| REMOVED | Drive has been physically removed | Reinsert or replace |
Common Pitfalls
Section titled “Common Pitfalls”Using RAIDZ1 with Large Drives
Section titled “Using RAIDZ1 with Large Drives”With modern drives (8 TB+), the probability of encountering an unrecoverable read error (URE) during A resilver approaches certainty. A RAIDZ1 pool with 8 TB drives has a resilver time of 12–24 hours. During that time, reading every block on every remaining drive means the chance of hitting a URE (and losing the pool) is non-trivial. Use RAIDZ2 (or RAIDZ3) for any pool with drives larger than 4 TB.
Setting dedup=on Without Sufficient RAM
Section titled “Setting dedup=on Without Sufficient RAM”Deduplication maintains an in-memory hash table of every unique block. This table requires Approximately 320 bytes per unique block. A 10 TB pool with 4 TB of unique data can require 100+ GB Of RAM for the dedup table. If the system runs out of RAM and must swap, performance collapses. Only Enable dedup if your data is highly redundant (VM templates, ISO images) and you have sufficient RAM. In most cases, compression (lz4) provides better space savings with no memory cost.
Mixing Drive Sizes in a RAIDZ Vdev
Section titled “Mixing Drive Sizes in a RAIDZ Vdev”While ZFS allows mixing drive sizes in a RAIDZ vdev, the pool capacity is determined by the smallest Drive in the vdev. A RAIDZ2 vdev with three 12 TB drives and one 4 TB drive will have the capacity Of four 4 TB drives. Always use identical drives within a vdev.
Not Setting ashift Correctly
Section titled “Not Setting ashift Correctly”Once a pool is created, ashift cannot be changed. Creating a pool with ashift=9 (512 bytes) on Drives with 4 KB physical sectors causes severe read-modify-write amplification on small writes, Reducing performance by 50–80%. Always use ashift=12 or higher.
Ignoring Fragmentation
Section titled “Ignoring Fragmentation”ZFS pools become fragmented over time due to the copy-on-write nature. Fragmentation above 70–80% Can significantly reduce performance, especially for random read workloads. Monitor fragmentation With zpool list -v. There is no native defragmentation tool for ZFS — the only way to defragment Is to copy the data to a new pool. Regular snapshot pruning and avoiding small random writes on HDD Pools help keep fragmentation manageable.
ZFS Pool Design Patterns
Section titled “ZFS Pool Design Patterns”Mirror Pool Design
Section titled “Mirror Pool Design”Mirrors are the gold standard for performance and redundancy:
# 2-way mirror (most common)zpool create -o ashift=12 -O compression=lz4 -O atime=off tank \ mirror /dev/sda /dev/sdb \ mirror /dev/sdc /dev/sdd \ mirror /dev/sde /dev/sdf
# Performance characteristics:# Read: N × single-disk IOPS (any disk in a mirror can serve the read)# Write: N × single-disk IOPS (writes go to all mirrors simultaneously)# Capacity: 50% of total raw# Fault tolerance: 1 disk per mirror vdevMirror pools provide the best random I/O performance because every vdev can serve reads Independently. A 6-disk mirror pool (3 mirror vdevs) can serve 3× the random IOPS of a single disk.
RAIDZ2 Pool Design
Section titled “RAIDZ2 Pool Design”RAIDZ2 provides dual-parity protection at better capacity efficiency:
# RAIDZ2 with 8 drives per vdevzpool create -o ashift=12 -O compression=lz4 -O atime=off tank \ raidz2 /dev/sda /dev/sdb /dev/sdc /dev/sdd /dev/sde /dev/sdf /dev/sdg /dev/sdh \ raidz2 /dev/sdi /dev/sdj /dev/sdk /dev/sdl /dev/sdm /dev/sdn /dev/sdo /dev/sdp
# Performance characteristics:# Read: Good (reads span all data disks)# Write: Moderate (parity calculation overhead)# Capacity: (N-2)/N of total raw per vdev# Fault tolerance: 2 disks per vdevdRAID (Declustered RAID)
Section titled “dRAID (Declustered RAID)”DRAID is a ZFS feature that distributes spare capacity across all drives in the pool, rather than Dedicating entire drives as hot spares:
# dRAID2 with distributed spareszpool create -o ashift=12 tank \ draid2:2d:8c:2s /dev/sda /dev/sdb /dev/sdc /dev/sdd /dev/sde /dev/sdf /dev/sdg /dev/sdh \ /dev/sdi /dev/sdj
# Parameters:# 2d = 2 data drives per stripe# 8c = 8 children (drives) per redundancy group# 2s = 2 distributed sparesDRAID provides faster resilvering than traditional RAIDZ because all drives participate in Rebuilding simultaneously.
Hybrid Pool Design (Special Vdev + Data Vdev)
Section titled “Hybrid Pool Design (Special Vdev + Data Vdev)”For workloads with mixed metadata and data requirements:
# NVMe metadata vdev + HDD data vdevzpool create -o ashift=12 tank \ mirror /dev/nvme0n1 /dev/nvme1n1 \ raidz2 /dev/sda /dev/sdb /dev/sdc /dev/sdd /dev/sde /dev/sdf
# Assign special small blocks to NVMezfs create -o special_small_blocks=32K tank/dataThis stores metadata (directories, file attributes) on the fast NVMe vdev while data resides on the HDD vdev, dramatically improving directory listing performance.
ZFS Compression Analysis
Section titled “ZFS Compression Analysis”Compression Ratio by Data Type
Section titled “Compression Ratio by Data Type”| Data Type | lz4 Ratio | zstd-3 Ratio | Compressible |
|---|---|---|---|
| Text files (source code, docs) | 2.0–3.0x | 2.5–4.0x | Yes |
| JSON, XML, CSV | 3.0–5.0x | 4.0–7.0x | Yes |
| Virtual machine images | 1.3–2.0x | 1.5–2.5x | Partially |
| Databases (relational) | 1.2–1.5x | 1.3–1.8x | Partially |
| Encrypted data | 1.0x | 1.0x | No |
| Media (JPEG, MP4, MKV) | 1.0x | 1.0x | No |
| Compressed archives (ZIP, tar.gz) | 1.0x | 1.0x | No |
| Logs (server, application) | 5.0–10.0x | 8.0–15.0x | Yes |
CPU Overhead of Compression
Section titled “CPU Overhead of Compression”| Algorithm | Compression Throughput | Decompression Throughput | CPU Overhead |
|---|---|---|---|
| lz4 | 3–5 GB/s per core | 8–12 GB/s per core | Minimal |
| zstd-1 | 1–2 GB/s per core | 4–6 GB/s per core | Low |
| zstd-3 | 500 MB–1 GB/s per core | 3–5 GB/s per core | Moderate |
| zstd-10 | 100–200 MB/s per core | 2–3 GB/s per core | High |
On modern CPUs (8+ cores), lz4 compression overhead is negligible for most workloads. The I/O time Saved by writing less data to disk exceeds the CPU time spent compressing.
When to Disable Compression
Section titled “When to Disable Compression”Disable compression only for data that is already compressed or encrypted:
# Disable compression for an existing datasetzfs set compression=off tank/media/movieszfs set compression=off tank/backups/encryptedzfs set compression=off tank/software/isosZFS ARC Internals
Section titled “ZFS ARC Internals”ARC Replacement Policy
Section titled “ARC Replacement Policy”The ARC maintains five lists:
- MRU (Most Recently Used): Ghost + active MRU lists.
- MFU (Most Frequently Used): Ghost + active MFU lists.
- Metadata (ARC meta): Separate cache for metadata (dnode structures, directory entries).
The replacement algorithm:
- New data enters the MRU list.
- On a cache hit, data is promoted from MRU to MFU (if accessed multiple times).
- On eviction, data moves from active to ghost list. Ghost entries remember the data’s identity but not its content.
- If a ghost entry is accessed again (cache miss → hit in ghost list), the data is fetched from disk and placed at the head of the appropriate active list.
- The ARC size is bounded by
arc_max(primary cache) andarc_meta_limit(metadata cache).
ARC Metadata Limit
Section titled “ARC Metadata Limit”Metadata (directory entries, file attributes, indirect blocks) can consume a significant portion of The ARC. The arc_meta_limit parameter controls the maximum fraction of ARC dedicated to metadata:
# Default: 1/4 of ARC# Recommended for metadata-heavy workloads: 1/2 to 3/4
# Check current metadata usagekstat -p zfs:0:arcstats:arc_meta_usedkstat -p zfs:0:arcstats:arc_meta_max
# The metadata-to-data ratio indicates workload characteristics:# High metadata/data ratio → Many small files (mail server, source code, home directories)# Low metadata/data ratio → Few large files (media, VM images, backups)ZFS Dataset Properties Reference
Section titled “ZFS Dataset Properties Reference”Compression Properties
Section titled “Compression Properties”| Property | Values | Default | Description |
|---|---|---|---|
| compression | off, lz4, lzjb, zstd, gzip-1..9, zle | lz4 (TrueNAS) | Compress data before writing |
| compressratio | Read-only | 1.00x | Current compression ratio |
Access Time Properties
Section titled “Access Time Properties”| Property | Values | Default | Description |
|---|---|---|---|
| atime | on, off | on | Update file access time on read |
| relatime | on, off | off | Update atime only if modified since last read |
| xattr | on, off | on | Enable extended attributes |
Quota and Reservation Properties
Section titled “Quota and Reservation Properties”| Property | Values | Default | Description |
|---|---|---|---|
| quota | Size or none | none | Maximum space for dataset + children |
| refquota | Size or none | none | Maximum space for dataset only (not children) |
| reservation | Size or none | none | Minimum space guaranteed for dataset |
| refreservation | Size or none | none | Minimum space for dataset only |
Sync Properties
Section titled “Sync Properties”| Property | Values | Default | Description |
|---|---|---|---|
| sync | standard, always, disabled | standard | Synchronous write behavior |
| logbias | latency, throughput | latency | Optimize for latency or throughput |
ZFS Snapshot Advanced Usage
Section titled “ZFS Snapshot Advanced Usage”Snapshot Naming Conventions
Section titled “Snapshot Naming Conventions”# Recommended naming convention:tank/data@daily-2024-01-15tank/data@weekly-2024-W03tank/data@monthly-2024-01tank/data@pre-upgrade-2024-01-15T10-30-00tank/data@manual-description
# List snapshots with sortingzfs list -t snapshot -s creation -o name,creation,used,referSnapshot Diff
Section titled “Snapshot Diff”# Show differences between two snapshotszfs diff tank/data@snap1 tank/data@snap2
# Output format:# M + path # Modified file# M - path # Deleted file# + + path # New file# R + old -> new # Renamed fileSnapshot Rollback
Section titled “Snapshot Rollback”# Rollback a dataset to a specific snapshot (DESTRUCTIVE: destroys all snapshots taken after)zfs rollback tank/data@daily-2024-01-10
# Force rollback (discard changes since snapshot)zfs rollback -rf tank/data@daily-2024-01-10ZFS Send/Receive Advanced Usage
Section titled “ZFS Send/Receive Advanced Usage”Raw Send for Encrypted Datasets
Section titled “Raw Send for Encrypted Datasets”# Raw send preserves encryption without requiring the key on the receiving sidezfs send -Rwv tank/encrypted@snap1 | ssh remote zfs recv -F backup/encrypted
# The receiving system cannot read the data without the encryption key# This is ideal for offsite backup where the remote system should not have accessResuming Interrupted Transfers
Section titled “Resuming Interrupted Transfers”# Send with resume token (saved periodically)zfs send -Rv -t <resume-token> | ssh remote zfs recv -s backup/data
# The resume token is printed when a transfer is interrupted (Ctrl+C)# Save it and use it to resume the transfer laterBandwidth-Limited Transfer
Section titled “Bandwidth-Limited Transfer”# Use pv to limit bandwidthzfs send -Rcv tank/data@snap1 | pv --rate-limit 50m | ssh remote zfs recv -F backup/data
# 50m = 50 MB/s# Adjust based on available bandwidth and impact on production workloadsZFS Pool Maintenance
Section titled “ZFS Pool Maintenance”Pool Expansion
Section titled “Pool Expansion”# Replace a smaller disk with a larger one (one at a time)zpool replace tank /dev/sda /dev/sdb-new
# After all disks in a vdev are replaced with larger disks:# The vdev automatically expands to use the full capacityzpool list -v
# Add a new vdev to the pool (stripes across vdevs)zpool add tank mirror /dev/sdg /dev/sdh
# Add a cache device (L2ARC)zpool add cache tank /dev/nvme0n1
# Add a log device (SLOG)zpool add log tank /dev/nvme1n1p1
# Add a spare devicezpool add spare tank /dev/sdiPool Export and Import
Section titled “Pool Export and Import”# Export a pool (unmount all datasets)zpool export tank
# Import a poolzpool import tank
# Import a pool from a specific cachefile (after disk replacement)zpool import -c /path/to/zpool.cache tank
# Import by GUID (more reliable than by name)zpool import <guid>
# Force import (if pool was not properly exported)zpool import -f tankCross-References
Section titled “Cross-References”- Sharing and Permissions — ZFS datasets are shared through protocols like SMB and NFS, connecting pool management to file access.
- Backup and Replication — ZFS snapshots are the foundation for replication-based backup strategies.
- ZFS Encryption — Encryption is configured at the dataset level, building on the ZFS pool and dataset architecture.