Skip to main content
Mini PC Lab logo
Mini PC LabMini PCs for Homelabs
guides

SSD Endurance for ZFS: TBW, SLOG, and Special Vdevs

By Max · October 2, 2026

SSD modules beside a compact ZFS NAS with mirrored storage groups

TBW is useful, but it is not a countdown clock. The number tells you how much host data a drive is rated to accept under a stated endurance model. Your ZFS workload can be lighter or heavier than that model, and the drive can show warning signs before the rating or continue operating beyond it without becoming a safe bet.

The right question is not simply whether an SSD has a high TBW number. Ask where the drive will live, what kind of writes it will receive, whether those writes are synchronous, how much ZFS work happens behind the scenes, and what happens if that device fails.

This guide turns those questions into a practical plan. For the interface choice first, read our NVMe versus SATA SSD mini NAS guide. For a broader TrueNAS hardware decision, use our TrueNAS mini PC guide.

TBW in plain English

TBW means terabytes written. It is a rating for the amount of host data written to an SSD over the endurance period defined by the manufacturer. The rating can vary by capacity, NAND type, workload class, temperature, spare area, and warranty terms.

DWPD means drive writes per day. A rating of one DWPD for five years means the drive is designed around one full drive capacity of host writes per day during that period. The exact definition depends on the product’s documentation and test conditions.

TBW and DWPD are not the same as a guarantee that the drive suddenly fails at the limit. They are also not a complete reliability score. A drive can encounter a controller fault, firmware issue, power-loss problem, or thermal problem before its media reaches the rated write total. A drive can also remain operational after the rated total while no longer being covered by its endurance warranty.

Manufacturers use different endurance classes for client and enterprise workloads. Seagate’s summary of the JEDEC JESD218 endurance standard shows how different those assumptions are. A client-class rating assumes 40 degrees Celsius for 8 hours a day of active use and must still retain data for one year powered off at 30 degrees. An enterprise-class rating assumes 55 degrees for 24 hours a day and three months of power-off retention at 40 degrees. A home NAS runs around the clock, which is closer to the enterprise pattern even when the drive is sold as a client model. Compare the workload assumptions, not just the largest number printed in a listing.

Endurance worksheet

Use this worksheet for each SSD role. Fill it in with your own numbers and the drive’s data sheet.

LineWhat to enterExample
A: rated enduranceTBW from the data sheet, or DWPD × capacity in TB × 365 × warranty years1 DWPD on a 2TB drive over 5 years is 1 × 2 × 365 × 5 = 3,650TB
B: headroomThe share of A you refuse to spend, for bursts, resilvers, and measurement error40 percent
C: planning budgetA × the remaining share3,650 × 0.6 = 2,190TB
D: daily host writesMeasured from the drive’s counters over a representative week, divided by 7250GB, or 0.25TB
E: annual host writesD × 36591TB
F: planning yearsC ÷ EAbout 24 years
G: check against warrantyCompare F with the warranty term; the shorter one is your planning horizon5 years

The same 250GB a day against a drive rated at 700TBW gives a very different answer. With 40 percent headroom the planning budget is 420TB, and at 91TB a year that lasts about 4.6 years, short of a five-year plan. That is the kind of result that should change which drive goes into a busy role.

None of these numbers is a promised lifespan. They compare your workload with the manufacturer’s rating model and tell you when to plan a replacement.

Measure line D on the real system after it has been running. Read the drive’s host-written counter at the start and end of a representative week, then repeat during a backup-heavy or migration-heavy month. A quiet week is not enough evidence for a server that runs large replication jobs every weekend.

Host writes are not the whole wear story

The drive’s counters report host writes, meaning what the operating system sent. The flash may do more work internally. Garbage collection, wear leveling, metadata updates, replacement blocks, and partially filled pages cause the drive to write more to its NAND than the host sent, which is called write amplification. JESD218 defines TBW in terms of host writes under a specified workload, so a workload that amplifies more than the rating workload will reach the NAND’s limits sooner than the host-write total suggests. The drive’s “Percentage Used” value, covered below, is the vendor’s own wear estimate, which makes it a better guide to remaining life than the host-write total alone.

This is why the same host-write total can produce different wear on two drives. A mostly sequential workload with free space and a controller that can consolidate writes may be gentler than a nearly full drive receiving small random updates. Compression can reduce physical data written in some workloads, while incompressible data does not receive that benefit.

ZFS adds its own behavior. Copy-on-write writes new blocks rather than overwriting old blocks in place. Snapshots preserve old blocks until they are released. Replication reads and writes large amounts of data. Scrubs read the pool, while resilvers can read and write heavily after a device failure. None of these activities makes ZFS unsafe, but they belong in the endurance estimate.

What creates writes in a homelab

Source of writesWhy it mattersHow to estimate it
File copies and downloadsLarge sequential host writesMeasure the weekly data added and replaced
DatabasesSmall random and synchronous patternsCheck database growth, logs, and churn
Virtual machinesGuest writes plus filesystem activityWatch virtual disk writes during normal and busy periods
Container dataLogs, indexes, thumbnails, and downloadsInclude application directories, not only final media
SnapshotsOld blocks remain until snapshots expireCount snapshot retention and changed data
ReplicationA second copy creates another write pathInclude both source and destination writes
ScrubsMostly reads, but can expose weak mediaTrack duration and errors rather than treating it as TBW
ResilversReads the surviving pool and writes replacement mediaReserve endurance headroom for a failed-device event
SLOGSynchronous write traffic reaches the log deviceMeasure sync-heavy workloads rather than all NAS traffic

The largest mistake is counting only the files you intentionally save. A photo library may be copied once but generate thumbnails, search indexes, previews, backups, and snapshots. A VM may hold a small operating system image while producing constant guest writes. Your endurance estimate needs the work created by the service, not only the size visible in the file browser.

ZFS roles do not stress an SSD the same way

An SSD used as a normal data vdev is part of the primary pool. An SSD used for L2ARC is a read cache that stores another copy of data already in the pool. An SSD used for SLOG receives the ZFS intent log for synchronous writes. An SSD used as a special vdev stores selected allocation classes as primary pool data.

ZFS roleWhat it storesEndurance questionFailure consequence
Data vdevNormal user dataHow much data and churn does the pool receiveDepends on the vdev redundancy
L2ARCA second copy of cached dataWill the cache receive enough reads and metadata activity to justify its writesLosing it should not lose pool data
SLOGIntent-log records for synchronous writesHow much sync traffic arrives and does the device have power-loss protectionOpenZFS says losing a modern SLOG costs at most the last few seconds of synchronous writes, not the pool; a device that lies about power loss can lose those writes silently
Special vdevMetadata, indirect blocks, and optional small blocksHow much metadata and small-block data will be placed thereLosing it loses the pool; on a raidz pool it can never be removed

The role changes the endurance calculation and the risk calculation. A high-TBW drive can still be the wrong device for a SLOG if it reports writes as stable before they are protected during a power loss. A mirrored special vdev can still be undersized if metadata and selected small blocks consume its headroom.

SLOG is not a general write cache

The ZFS Intent Log exists to satisfy synchronous write semantics. A dedicated SLOG moves that log to a separate device. It does not turn every asynchronous file copy into a faster write, and it does not make a slow pool universally faster.

OpenZFS states that asynchronous writes do not touch the ZIL and that a SLOG is read after a crash to replay work that had not yet been committed. The workloads that can benefit include some NFS, database, and virtual machine patterns that issue synchronous writes. If your NAS mainly serves media files over ordinary SMB copies, a SLOG may add complexity without solving the real limit.

Power-loss protection matters because the device must honor the promise that a completed synchronous write is on stable storage. A consumer SSD may have a high sequential speed and an attractive TBW rating while lacking the capacitors and firmware behavior needed for this role. OpenZFS recommends a low-latency device with power-loss protection for SLOG use.

Before adding one, verify that the workload issues synchronous writes, measure where the delay is, and confirm the drive’s power-loss behavior from a technical data sheet. Do not use sync=disabled as a speed trick for data that must survive a power failure. That setting trades durability for performance.

Special vdev is permanent pool storage

OpenZFS describes a special vdev as a top-level vdev for metadata, indirect blocks, deduplication tables, and optional small file blocks. Those blocks live there rather than as a second copy in the normal class. That is why the documentation warns that the special vdev must be at least as redundant as the normal vdevs.

This has two practical consequences.

First, do not add a single spare NVMe drive as a special vdev to an otherwise redundant pool. The pool may appear healthy until that one device fails, at which point the special allocation class can take the pool with it.

Second, do not size it as if it were a temporary cache. OpenZFS says allocations that no longer fit spill back to the normal class, existing blocks are never migrated in either direction, and by default the last quarter of the special class is reserved for metadata alone. Leave headroom and monitor the class with zpool list -v.

Third, treat adding one as permanent. OpenZFS states that on a raidz pool a special vdev can never be removed. On a mirror-only pool removal is possible, but only under the same constraints as any top-level vdev removal. If you are unsure, do not add it.

A mirrored special vdev can be useful when metadata-heavy work on a disk pool is the real problem. It is not a shortcut to turn a weak pool layout into a fast all-flash pool, and it is not a substitute for a proper backup.

L2ARC has a different trade-off

L2ARC is a secondary read cache. It stores data that also exists in the main pool, so losing the cache should not lose the primary data. That makes its failure consequence gentler than a special vdev failure.

The trade-off is that L2ARC consumes device space, RAM for cache metadata, and SSD write endurance as the cache is populated and managed. OpenZFS warns that on a RAM-constrained system, L2ARC can make things slower. Add it only after more RAM and a correctly sized primary pool have been considered.

For most small homelabs, a faster data pool, more memory, or a better dataset layout is a clearer first step than adding several cache devices. A cache is useful when it solves a measured repeated-read problem, not because an empty M.2 slot exists.

How to choose an endurance class

Use the role and workload together.

Ordinary data vdev

Choose a drive whose endurance rating covers the estimated host writes with useful headroom. Look for reliable health reporting, a documented temperature range, and a warranty that explains how endurance limits apply. Redundancy belongs in the vdev layout, not in the TBW number.

Read cache

Prioritize adequate capacity, health monitoring, and low enough cost that replacing the device is easy. Do not use a low-end drive for a cache that constantly churns if its endurance is marginal. If the workload does not benefit from the cache, remove the complexity instead.

SLOG

Treat power-loss protection and latency as requirements. Then check endurance against the synchronous write rate. A SLOG does not need to hold the whole pool, but it does need enough space for the workload and a design that can survive a device failure without violating the pool’s redundancy plan.

Special vdev

Treat the device as primary pool storage. Match its redundancy to the normal pool, leave generous space for metadata growth, and avoid a layout that cannot be replaced or removed later. The endurance estimate should include metadata churn, snapshots, small blocks, and the administrative work created by the pool.

Monitor the drive after deployment

A printed TBW number is only useful if you compare it with the live device state. Track these signals:

  • Host data units written or the equivalent counter exposed by the drive
  • Percentage used or endurance remaining
  • Available spare capacity
  • Critical warnings and media errors
  • Composite temperature and thermal events
  • ZFS read, write, checksum, and device errors
  • Pool capacity and special-class usage
  • Snapshot and replication growth
  • Resilver duration and replacement history

For an NVMe drive, the standard SMART and health log carries most of these. Read it with smartmontools or nvme-cli. Both commands are read-only and need root:

# NVMe health report
smartctl -a /dev/nvme0

# Same log via nvme-cli 3.x
# older nvme-cli: nvme smart-log
nvme log smart /dev/nvme0

# ZFS errors and per-vdev use
zpool status -v
zpool list -v

What to read and when to act:

FieldWhat it meansAct when
Critical WarningBit flags for spare below threshold, temperature out of range, reliability degraded, media read-only, and failed volatile-memory backupAny value other than zero
Available Spare and its thresholdRemaining spare capacity as a percentageSpare approaches the threshold
Percentage UsedThe vendor’s estimate of endurance consumedIt crosses the replacement point you set, for example 80 percent
Data Units WrittenHost writes, counted in units of 1,000 × 512 bytes; smartctl prints the byte total beside itUse it for line D of the worksheet
Media and Data Integrity ErrorsUnrecovered data errorsThe count rises
Unsafe ShutdownsPower lost without a clean shutdownIt rises on a SLOG device, which puts power-loss protection to the test

The field names and flag meanings above are those smartmontools prints for the NVMe health log. SATA SSDs report vendor-specific attributes instead, so check the drive maker’s documentation for the equivalent wear and spare values. The 80 percent replacement point is our planning suggestion, not a vendor rule.

Set an alert before the first device becomes critical. Keep a replacement drive or a realistic procurement path for the pool. A mirrored pool can keep serving data after one device fails, but redundancy buys time to replace the drive. It does not make the failed device healthy or remove the need for a tested backup.

Why heat and free space matter

Endurance is not independent of the enclosure. A hot SSD may throttle, spend more time doing background work, or become less predictable under sustained writes. M.2 devices in dense mini NAS systems need a real thermal plan rather than a decorative cover.

Free space also matters. A nearly full SSD gives the controller fewer places to consolidate data and perform wear leveling. Leave operational headroom in the pool and in individual datasets. Do not confuse unused ZFS capacity with an SSD vendor’s reserved area, because they serve different purposes.

Compression can reduce physical writes when the data compresses well. It cannot be assumed for encrypted or already compressed media. Measure the actual workload and use the pool’s statistics rather than applying one universal write multiplier.

Who should skip a SLOG or special vdev

Skip a SLOG when the NAS mainly serves asynchronous SMB files, media, photos, backups, and documents. Spend the effort on a proper data layout, network path, UPS, and backup instead.

Skip a special vdev when you cannot provide redundant devices, monitor its capacity, or accept the pool-layout constraints that can make removal difficult. A single special device is not a harmless experiment.

Skip L2ARC when the system is short on RAM or you have not identified a repeated-read workload that misses in ARC. More cache hardware is not a replacement for enough memory.

Skip an endurance calculation based only on the drive’s marketing number. If you do not know the write workload, start with a measured baseline and revisit the decision after several weeks.

A practical endurance checklist

Before assigning an SSD to a ZFS role:

  1. Estimate daily host writes from applications, users, backups, snapshots, and replication.
  2. Measure the live host-write counter during both quiet and busy periods.
  3. Compare annual writes with the vendor’s TBW or DWPD model and keep headroom.
  4. Check whether the workload is mostly asynchronous, synchronous, sequential, or random.
  5. Confirm temperature, health counters, and power-loss behavior.
  6. Match vdev redundancy to the consequence of losing the device.
  7. Decide how the device will be replaced before adding it to the pool.
  8. Confirm that a separate backup can restore the data if the whole pool is lost.

The NAS capacity calculator helps with the capacity and redundancy side of the decision. Our mini PC home server backup guide covers the separate protection layer that endurance ratings cannot provide.

Final answer

TBW is a planning input, not a promise. Estimate writes, compare the result with the manufacturer’s workload model, leave headroom, and monitor the actual drive after deployment.

For ordinary ZFS data storage, choose a drive with an endurance rating that fits the writes and a vdev layout that survives the failure you can afford. Add L2ARC only for a measured read-cache problem. Add a SLOG only for synchronous writes and only with power-loss protection. Add a special vdev only when you can mirror it and accept that its blocks are primary pool data.

The safest ZFS SSD is not always the one with the biggest TBW number. It is the one whose workload, protection, cooling, monitoring, and replacement plan all agree.


Frequently Asked Questions

What does TBW mean for an SSD?

TBW means terabytes written. It is an endurance rating for the amount of host data a drive is expected to write under a stated workload and warranty model. It is useful for comparison, but it is not a promise that the drive fails at that exact number.

Is a consumer SSD safe for ZFS?

It can be fine as an ordinary data device or read cache when its endurance and temperature fit the workload. A consumer SSD without power-loss protection is a poor choice for a SLOG because synchronous writes depend on stable storage.

Does ZFS wear out SSDs faster?

ZFS can create more internal work through copy-on-write, snapshots, replication, scrubs, resilvers, and small random writes. The effect depends on the workload, pool layout, compression, free space, and write amplification rather than on ZFS alone.

Does a special vdev need to be mirrored?

Yes. A special vdev stores metadata and optional small blocks that exist only there. OpenZFS says it must be at least as redundant as the pool’s normal vdevs because losing it can lose the pool.

Do I need a SLOG in a home NAS?

Usually not. A SLOG helps only synchronous write workloads such as some NFS, database, and virtual machine patterns. It does not speed ordinary asynchronous file copies, and it needs a low-latency device with power-loss protection.


Sources and further reading