Articles & Snippets

How to Replace a Failed ZFS RAIDZ1 Drive on TrueNAS with an HP Smart Array P400i

This guide explains how to replace a failed drive in a TrueNAS ZFS RAIDZ1 pool when the system uses an HP Smart Array P400i controller.

1. Check the ZFS pool

Start by checking the pool:

sudo zpool status -v pool

Look for a drive showing FAULTED or UNAVAIL. In our case, the failed ZFS member was:

134bd9d1-d903-424c-a560-887f0c640bb4

The pool reported:

errors: No known data errors

This is important. It means ZFS did not report known corrupted data at that point.

2. Identify the physical drive behind the HP Smart Array controller

With an HP Smart Array P400i, Linux may not expose the physical disks directly. The controller can present logical volumes such as:

LOGICAL VOLUME  900G

To inspect the physical drives, first identify the controller's SCSI generic device. In our system it was /dev/sg6.

Then check the physical drives:

sudo smartctl -a -d cciss,0 /dev/sg6
sudo smartctl -a -d cciss,1 /dev/sg6
sudo smartctl -a -d cciss,2 /dev/sg6
sudo smartctl -a -d cciss,3 /dev/sg6
sudo smartctl -a -d cciss,4 /dev/sg6
sudo smartctl -a -d cciss,5 /dev/sg6

Match the serial number and model reported by SMART with the physical drive you removed from the server.

3. Replace the physical hard drive

Remove the failed physical drive and install the replacement drive. Make sure the replacement is compatible with the controller and is large enough for the existing logical volume.

After installing the replacement, allow the HP Smart Array controller to detect it. A reboot may be necessary depending on the controller and its current state.

4. Check the logical volumes after reboot

Run:

lsblk -o NAME,SIZE,MODEL,SERIAL,PARTUUID

In our case, the new logical volume appeared as:

sde  838.3G  LOGICAL VOLUME

But initially there was no sde1.

This is important because the existing RAIDZ members were using partitions, for example:

sdb1
sdc1
sdd1
sdf1

5. Do not run zpool replace yet

If /dev/sde1 does not exist, this command will fail:

sudo zpool replace pool 17215090055309476533 /dev/sde1

You must first create the replacement partition.

6. Copy the partition layout from a healthy RAIDZ member

We used the healthy /dev/sdf device as the partition-layout template. This copies the GPT partition structure only. It does not copy the data.

sudo sgdisk -R=/dev/sde /dev/sdf

Then generate new unique GUIDs for the replacement disk:

sudo sgdisk -G /dev/sde

Ask Linux to reread the partition table:

sudo partprobe /dev/sde

7. Confirm that the new partition exists

lsblk -o NAME,SIZE,MODEL,SERIAL,PARTUUID

You should now see something similar to:

sde
└─sde1  836.3G  ...  8b748d1e-b5b9-49e8-8ed8-f702ec553a96

8. Replace the failed ZFS member

Once /dev/sde1 exists, run:

sudo zpool replace pool 17215090055309476533 /dev/sde1

The number 17215090055309476533 represents the unavailable ZFS member from this particular example. Do not copy this number into another system.

Get the correct failed device identifier from your own:

sudo zpool status -v pool

9. Confirm that the resilver has started

sudo zpool status -v pool

You should see something similar to:

status: One or more devices is currently being resilvered.

replacing-3
  old-device  UNAVAIL
  sde1         ONLINE  (resilvering)

ZFS calls this process a resilver. It is essentially the process of reconstructing the missing RAIDZ data onto the replacement disk.

10. Let the resilver finish

Continue checking the status with:

sudo zpool status -v pool

The pool will normally remain DEGRADED while the resilver is running. That is expected.

Do not reboot the server or remove the replacement drive while the resilver is running.

The time required depends on the amount of data and the performance of the disks and controller. The initial estimated completion time can change as the resilver progresses.

11. Confirm the replacement is complete

When the resilver finishes, run:

sudo zpool status -v pool

The final result should show:

state: ONLINE

The temporary replacing-3 section should be gone, and the replacement drive should appear as a normal ONLINE member.

Also verify:

errors: No known data errors

Quick Replacement Checklist

  1. Check the ZFS pool with zpool status -v pool.
  2. Identify the failed physical disk by serial number.
  3. Replace the physical disk.
  4. Reboot if necessary for the HP Smart Array controller to detect it.
  5. Run lsblk -o NAME,SIZE,MODEL,SERIAL,PARTUUID.
  6. Confirm the new logical volume.
  7. If the replacement partition does not exist, copy the partition layout.
  8. Run sgdisk -R.
  9. Run sgdisk -G.
  10. Run partprobe.
  11. Confirm /dev/sde1 exists.
  12. Run zpool replace using the correct failed ZFS device.
  13. Monitor the resilver with zpool status -v pool.
  14. Wait until the pool returns to ONLINE.
  15. Confirm errors: No known data errors.

TrueNas  HP