UNSOLVED

dactylos

updated

7 years ago

D

dactylos

7 Posts

0

6036

April 1st, 2019 08:00

PS6000 RAID 6 Multiple disk failure

Hello all,

I'd really appreciate some feedback on an out of warranty Equallogic Array we are trying to recover data from.

I have a little experience with Equallogic arrays but at present any clue would be helpful and yes, I know the system is EOL out of warranty and should not contain productive data but that is the current situation and all we're trying to accomplish is to move the data that is still somehow available away and off the group.

To start off a bit of circumstantial info. The group had a complete failure last week and when looking at it, it became clear that there were multiple drive failures on member 2 which caused it to go offline and when checking the CLI it shows it was no longer initializing as the raid was in a failed and unrecoverable state thus prompting the remaining member to also take the group offline due to the missing member.

In the meanwhile we attempted to fit a drive in the array that could replace disk 7 that had previously completely failed and does no longer spin up.

The disk(s) in question were sourced from another EQL array we set back to factory defaults to clear all info and re-initialize the disks.

After replacing disk 7 the raid now shows it is in a degraded state with dirty raid cache and that the newly fitted disk is available as hotspare but no rebuilt has set in. We did try cold reboots, restarts form the CLI and a controller failover but there has been no change.

I may simply be looking at the wrong end of this so if anyone has an idea on how to proceed to get the rebuilt started that would be great. I'm also lookign to find out if there's any chance to recover partial data if the other healthy member is somehow made to set the LUN's or even better their snapshots online again (and yes I am well aware that the missing blocks will make it likely a corrupted mess, but we have a RDM with flat file content that could possibly be all we urgently want back as we have at least partial backups).

Sorry for the wall of text but I thought I should try and be descriptive.

Below the more technical details.

Current status of Equallogic Group

Members: 2
Firmware: 8.1.4
Member 1: PS4000 RAID 5 Healthy
Member 2: PS6000 RAID 6 degraded due to 2 Failed drives (1/13), Disk 7 assigned as hotspare but not rebuilding. Disks with SMART trip alert: 1/4/5/13

CLI Outputs:

CLI(support)> raidtool
Driver Status: Ok
RAID LUN 0 Degraded.
raid status dirty.
15 Drives (0,14,4,6,8,10,12,f,3,5,2,9,15,f,11)
RAID 6 (64KB sectPerSU)
Capacity 7,489,541,636,096 bytes
Available Drives List: 7
Unavailable Drives:
1 (history of failure)
13 (history of failure)
CLI(support)> raidtool -Z
Active RAID LUNs: 0
Driver Status = driver running.
Malloc Bytes = 0KB
Outstanding Active I/O's = 0
Pending I/O's = 0
Pending Resource Reqs = 0
Outstanding StripeLocks = 0
Allocated Sectors = 0

Device = 000
status = 002
outio = 00000000
drives = 13
disk luns: 11 15 9 2 5 3 12 10 8 6 4 14 0
disk count:16
disk lun= 0 status=0x00000400 drive active device=0
disk lun= 1 status=0x00020000 history of failure no-device
disk lun= 2 status=0x00000400 drive active device=0
disk lun= 3 status=0x00000400 drive active device=0
disk lun= 4 status=0x00000400 drive active device=0
disk lun= 5 status=0x00000400 drive active device=0
disk lun= 6 status=0x00000400 drive active device=0
disk lun= 7 status=0x00001000 hot spare no-device
disk lun= 8 status=0x00000400 drive active device=0
disk lun= 9 status=0x00000400 drive active device=0
disk lun=10 status=0x00000400 drive active device=0
disk lun=11 status=0x00000400 drive active device=0
disk lun=12 status=0x00000400 drive active device=0
disk lun=13 status=0x00020000 history of failure no-device
disk lun=14 status=0x00000400 drive active device=0
disk lun=15 status=0x00000400 drive active device=0
CLI(support)> exec "raidtool -w 0"
opendisk failed 19 Operation not supported by device
CLI(support)>