ZFS Pool failure

I have a faulty disk in raidz1 pool, it keeps disappearing from time to time. So, I removed it and replaced. The funny situation that the new disk was faulty as well.

Resilvering started, but emitted checksum errors. So, I shutdowned the system and returned faulty drive. I've planned to online it and resilver.
The problem:

Code:
abishai@beta:~ % doas zpool import
  pool: zdata
    id: 6286655327721895903
 state: FAULTED
status: The pool metadata is corrupted.
action: The pool cannot be imported due to damaged devices or data.
    The pool may be active on another system, but can be imported using
    the '-f' flag.
   see: https://openzfs.github.io/openzfs-docs/msg/ZFS-8000-72
config:

    zdata                                             FAULTED  corrupted data
      raidz1-0                                        DEGRADED
        replacing-0                                   DEGRADED
          gptid/5416f821-9314-11ef-b60b-a8b8e000d71b  ONLINE
          gpt/data0                                   FAULTED  corrupted data
        gpt/data1                                     ONLINE
        gpt/data2                                     ONLINE
        gpt/data3                                     ONLINE
    logs
      mirror-1                                        ONLINE
        gpt/zil0                                      ONLINE
        gpt/zil1                                      ONLINE

1. data0 is faulty resilvering target, removed.
2. gptid/5416f821-9314-11ef-b60b-a8b8e000d71b is the disk that likes to disappear. It represents pool state in the past (say, 1 day ago, because it wasn't resilvered after disappearing to the present state).

Other disks are OK.

Code:
abishai@beta:~ % doas zpool import -F zdata
cannot import 'zdata': I/O error
    Destroy and re-create the pool from
    a backup source.

The pool shows problem with 1 disk, from my point of view it has enough replicas to function.
I've emitted -FX, it didn't returned, but the disks show activity. Probably, I shouldn't use -X before asking for help and definitely should try to import without both data0 and gptid/5416f821-9314-11ef-b60b-a8b8e000d71b , but still - why failure in one device brought entire pool down?
 
Note it suggested -f, not -F. I would suggest removing the flaky drive and run on the three good drives until you (soon, hopefully) have a healthy replacement.

I believe it is currently confused because the last-update timestamps on the removed-and-returned flaky drive don’t line up with the other three. -f may let you get back up and running with the flaky drive present, but if it’s blinking in and out, why bother?
 
Note it suggested -f, not -F. I would suggest removing the flaky drive and run on the three good drives until you (soon, hopefully) have a healthy replacement.

I believe it is currently confused because the last-update timestamps on the removed-and-returned flaky drive don’t line up with the other three. -f may let you get back up and running with the flaky drive present, but if it’s blinking in and out, why bother?
I've tried -f, didn't helped. The primary issue is FAULT state. My mistake, I should really try to force-import pool without both blinking and faulty drives.

Now it still executing zpool import -FX zdata. 3 drives of pool are 100% active according to gstat. Blinking drive is not used by zpool at all. I think is is unwise to abort operation. No idea what it's try to do. From my perpective, 3 drives are intact, 1 in the past (somehow it is detected as online - I believe it shouldn't)

After drive blinked the first time, I flipped it with another to be sure it is a drive itself, not cable or backplane. It wasn't confused, resilvering was very fast, as if it looked at timestamp and applied only fraction of data. I scrubbed pool after that. After it blinked from new position, I've tried to replace it and made situation even worse :/
 
I've tried -f, didn't helped. The primary issue is FAULT state. My mistake, I should really try to force-import pool without both blinking and faulty drives.

Now it still hanging after zpool import -FX zdata. 3 drives of pool are 100% active according to gstat. Blinking drive is not used by zpool at all. I think is is unwise to abort operation.

If you get hangs import readonly. `zpool import -o readonly=on <poolname>`.

The operations that make ZFS hang are usually the block allocation ones, and you don't have those in readonly mode.
 
  • Thanks
Reactions: mro
zpool import -FX zdata worked and the pool has returned back. No output was printed in console, nor in logs.

Current pool state:
Code:
abishai@beta:~ % zpool status zdata
  pool: zdata
 state: DEGRADED
status: One or more devices has experienced an unrecoverable error.  An
    attempt was made to correct the error.  Applications are unaffected.
action: Determine if the device needs to be replaced, and clear the errors
    using 'zpool clear' or replace the device with 'zpool replace'.
   see: https://openzfs.github.io/openzfs-docs/msg/ZFS-8000-9P
  scan: resilvered 0B in 07:51:46 with 0 errors on Mon Oct  5 06:23:35 2026
config:

    NAME                        STATE     READ WRITE CKSUM
    zdata                       DEGRADED     0     0     0
      raidz1-0                  DEGRADED     0     0     0
        replacing-0             UNAVAIL      0     0     0  insufficient replicas
          16309429118673301514  UNAVAIL      0     0     0  was /dev/gptid/5416f821-9314-11ef-b60b-a8b8e000d71b
          14546578228573414211  UNAVAIL      0     0     0  was /dev/gpt/data0
        gpt/data1               ONLINE       0     0     2
        gpt/data2               ONLINE       0     0    20
        gpt/data3               ONLINE       0     0    19
    logs
      mirror-1                  ONLINE       0     0     0
        gpt/zil0                ONLINE       0     0     0
        gpt/zil1                ONLINE       0     0     0

errors: No known data errors

Only the blinking drives walked out and in once.

Should I continue to use the pool after I'll receive replacement or full recreation is nesessary? Checksum errors are confusing, if there are not enough replicas, this means permanent damage, but no entries was printed.
 
Back
Top