ZFS Pool failure

I have a faulty disk in raidz1 pool, it keeps disappearing from time to time. So, I removed it and replaced. The funny situation that the new disk was faulty as well.

Resilvering started, but emitted checksum errors. So, I shutdowned the system and returned faulty drive. I've planned to online it and resilver.
The problem:

Code:
abishai@beta:~ % doas zpool import
  pool: zdata
    id: 6286655327721895903
 state: FAULTED
status: The pool metadata is corrupted.
action: The pool cannot be imported due to damaged devices or data.
    The pool may be active on another system, but can be imported using
    the '-f' flag.
   see: https://openzfs.github.io/openzfs-docs/msg/ZFS-8000-72
config:

    zdata                                             FAULTED  corrupted data
      raidz1-0                                        DEGRADED
        replacing-0                                   DEGRADED
          gptid/5416f821-9314-11ef-b60b-a8b8e000d71b  ONLINE
          gpt/data0                                   FAULTED  corrupted data
        gpt/data1                                     ONLINE
        gpt/data2                                     ONLINE
        gpt/data3                                     ONLINE
    logs
      mirror-1                                        ONLINE
        gpt/zil0                                      ONLINE
        gpt/zil1                                      ONLINE

1. data0 is faulty resilvering target, removed.
2. gptid/5416f821-9314-11ef-b60b-a8b8e000d71b is the disk that likes to disappear. It represents pool state in the past (say, 1 day ago, because it wasn't resilvered after disappearing to the present state).

Other disks are OK.

Code:
abishai@beta:~ % doas zpool import -F zdata
cannot import 'zdata': I/O error
    Destroy and re-create the pool from
    a backup source.

The pool shows problem with 1 disk, from my point of view it has enough replicas to function.
I've emitted -FX, it didn't returned, but the disks show activity. Probably, I shouldn't use -X before asking for help and definitely should try to import without both data0 and gptid/5416f821-9314-11ef-b60b-a8b8e000d71b , but still - why failure in one device brought entire pool down?
 
Note it suggested -f, not -F. I would suggest removing the flaky drive and run on the three good drives until you (soon, hopefully) have a healthy replacement.

I believe it is currently confused because the last-update timestamps on the removed-and-returned flaky drive don’t line up with the other three. -f may let you get back up and running with the flaky drive present, but if it’s blinking in and out, why bother?
 
Note it suggested -f, not -F. I would suggest removing the flaky drive and run on the three good drives until you (soon, hopefully) have a healthy replacement.

I believe it is currently confused because the last-update timestamps on the removed-and-returned flaky drive don’t line up with the other three. -f may let you get back up and running with the flaky drive present, but if it’s blinking in and out, why bother?
I've tried -f, didn't helped. The primary issue is FAULT state. My mistake, I should really try to force-import pool without both blinking and faulty drives.

Now it still executing zpool import -FX zdata. 3 drives of pool are 100% active according to gstat. Blinking drive is not used by zpool at all. I think is is unwise to abort operation. No idea what it's try to do. From my perpective, 3 drives are intact, 1 in the past (somehow it is detected as online - I believe it shouldn't)

After drive blinked the first time, I flipped it with another to be sure it is a drive itself, not cable or backplane. It wasn't confused, resilvering was very fast, as if it looked at timestamp and applied only fraction of data. I scrubbed pool after that. After it blinked from new position, I've tried to replace it and made situation even worse :/
 
I've tried -f, didn't helped. The primary issue is FAULT state. My mistake, I should really try to force-import pool without both blinking and faulty drives.

Now it still hanging after zpool import -FX zdata. 3 drives of pool are 100% active according to gstat. Blinking drive is not used by zpool at all. I think is is unwise to abort operation.

If you get hangs import readonly. `zpool import -o readonly=on <poolname>`.

The operations that make ZFS hang are usually the block allocation ones, and you don't have those in readonly mode.
 
  • Thanks
Reactions: mro
zpool import -FX zdata worked and the pool has returned back. No output was printed in console, nor in logs.

Current pool state:
Code:
abishai@beta:~ % zpool status zdata
  pool: zdata
 state: DEGRADED
status: One or more devices has experienced an unrecoverable error.  An
    attempt was made to correct the error.  Applications are unaffected.
action: Determine if the device needs to be replaced, and clear the errors
    using 'zpool clear' or replace the device with 'zpool replace'.
   see: https://openzfs.github.io/openzfs-docs/msg/ZFS-8000-9P
  scan: resilvered 0B in 07:51:46 with 0 errors on Mon Oct  5 06:23:35 2026
config:

    NAME                        STATE     READ WRITE CKSUM
    zdata                       DEGRADED     0     0     0
      raidz1-0                  DEGRADED     0     0     0
        replacing-0             UNAVAIL      0     0     0  insufficient replicas
          16309429118673301514  UNAVAIL      0     0     0  was /dev/gptid/5416f821-9314-11ef-b60b-a8b8e000d71b
          14546578228573414211  UNAVAIL      0     0     0  was /dev/gpt/data0
        gpt/data1               ONLINE       0     0     2
        gpt/data2               ONLINE       0     0    20
        gpt/data3               ONLINE       0     0    19
    logs
      mirror-1                  ONLINE       0     0     0
        gpt/zil0                ONLINE       0     0     0
        gpt/zil1                ONLINE       0     0     0

errors: No known data errors

Only the blinking drives walked out and in once.

Should I continue to use the pool after I'll receive replacement or full recreation is nesessary? Checksum errors are confusing, if there are not enough replicas, this means permanent damage, but no entries was printed.
 
The pool refuses to resilver. When I import clean disk, the process stops with a lot of read/write/checksum errors that leads to pool suspension. The pool is very old, 4x1Tb disks with > 70k hours. Probably, I should replace all of them with 2 bigger disks and not try to recover.

The readings are so confusing. I removed new data0 after pool suspension and rebooted, but zpool shows it's ... resilvering ?I wonder to know how..

When a device is replaced, a resilvering operation is initiated to movedata from the good copies to the new device. This action is a form of diskscrubbing. Therefore, only one such action can occur at a given time in thepool. If a scrubbing operation is in progress, a resilvering operation suspendsthe current scrubbing and restarts it after the resilvering is completed.

3 data errors has appeared, however -v shows no files.

Code:
abishai@beta:~ % zpool status
  pool: zdata
 state: DEGRADED
status: One or more devices is currently being resilvered.  The pool will
    continue to function, possibly in a degraded state.
action: Wait for the resilver to complete.
  scan: resilver in progress since Thu Oct  8 10:23:17 2026
    508G / 2.18T scanned at 457M/s, 140G / 2.18T issued at 126M/s
    0B resilvered, 6.29% done, no estimated completion time
config:

    NAME                        STATE     READ WRITE CKSUM
    zdata                       DEGRADED     0     0     0
      raidz1-0                  DEGRADED     0     0     0
        replacing-0             UNAVAIL      0     0     0  insufficient replicas
          16309429118673301514  UNAVAIL      0     0     0  was /dev/gptid/5416f821-9314-11ef-b60b-a8b8e000d71b
          14344583657725810227  UNAVAIL      0     0     0  was /dev/gpt/data0
        gpt/data1               ONLINE       0     0     0
        gpt/data2               ONLINE       0     0     0
        gpt/data3               ONLINE       0     0     0
    logs
      mirror-1                  ONLINE       0     0     0
        gpt/zil0                ONLINE       0     0     0
        gpt/zil1                ONLINE       0     0     0

errors: 3 data errors, use '-v' for a list
 
OK, I replicated initial problem, because I ended with the same non-importing pool and trying to import it with -FX once more (I hope it saves me one's more).
I've verified spare with write and read - it is working.
After I inserted spare and started resilvering another drive detaches almost immediately. Spare listed as 'ONLINE' and pool failure is not detected (resilver emits CHSUM errors for every block).

After that, pool is in FAULT state. -FX brings it online it looks like with some minor damage (volatile opened files, like logs, are broken ). Probably, because detached drive is slightly out of sync and -FX determines txd from it So, the question is if it is another drive failing or if system suddenly can't accommodate 4 drives.

This is definitely ZFS issue - pool must be shutdowned - no redundancy available, the spare shouldn't count before resilvering is complete.

Probably, I need stop experiments and backup the pool. And I don't have enough free disk space for all data.
 
Back
Top