UNSOLVED

UCTpro5

updated

12 years ago

U

UCTpro5

4 Posts

0

28771

August 4th, 2014 12:00

Perc5 & MD1000 lost RAID config

We have a PE2950 with a Perc 5 controller connected to two MD1000 units, each containing 12 500Gb SATA disks.

Last week, we had a power "blip". This caused one of the disks to die; it has been replaced.

After the "blip", we have also lost our entire RAID configuration. One MD1000 unit reports that all physical disks are "Ready" and the other reports that two disks are "Online" and the rest are "Ready". According to the controller, there is only one RAID-5 virtual disk set up consisting of the two "Online" disks and three others. This is clearly wrong.

We had thought to retag (but not initialize) the disks, but the system was installed in 2007 and we don't know the exact way RAID arrays were set up (all we know is that there was 7.2TB of usable space total). It is possible that each MD1000 has two RAID5 with a hot spare or one RAID5 with a hot spare (or something else along those lines). Is there any way to find out the original configuration? [We have run the DSET but it provided nothing.]

Any idea why most of the disks are reporting "Online" rather than "Foreign"? Any suggestions on best way to attempt recovery? We understand that each disk should have information about the RAID array it is in - where is this information stored on the disk?

  • DELL-Sam L

    Community Manager

    8058 Posts

    34386 Points

    632

    0

    Posted August 5th, 2014 06:00

    Hello UCTpro5,

    First off are your MD1000 daisy chained together or are they separate?  Second when you go into the PERC5 bios does it show the same information as OMSA does about the raid configuration or does it show something different?  Also while you are in the PERC5 bios check to see if the drive that you replaced is online & finished rebuilding & if your hotspare is still available.  

    Since you state that you don’t know how the raid configuration was setup prior to this event you may need to do a retag as the virtual disk info should still be on the drives, but would only do that if the PERC5 bios state the same as OMSA.

    Please let us know if you have any other questions.

  • DELL-Sam L

    Community Manager

    8058 Posts

    34386 Points

    632

    0

    Posted August 5th, 2014 10:00

    Hello UCTpro5,

    What I would suggest is that if you haven’t already contacted support I would open a ticket so that they can try a few things to see they can recover the raid info & assist with the retag in the correct order & save the data.  Other than that the other option that I see is that you may need to contact a data recovery company to see if they can cover the data if it is mission critical.

    Please let us know if you have any other questions.

  • UCTpro5

    4 Posts

    632

    0

    Posted August 5th, 2014 07:00

    Thanks for your response.

    They are separate. Yes, the PERC5 bios shows the same information as the OMSA (doesn't the OMSA just poll the bios?)

    The drive we replaced is in the same state “Ready” as 22 of the 24 drives. It cannot rebuild or act as a hot spare because the RAID configuration has been lost so the controller doesn’t know to assign it to a role.

  • UCTpro5

    4 Posts

    632

    0

    Posted August 5th, 2014 09:00

    An update on this: we have booted the PE server to a live distribution and accessed the server logs and configuration (we couldn't boot the server into its own OS). It seems that each of the MD1000s are split into two 5 disk RAID5 arrays (presumably) with two hot spares.

    I guess we will never be able to work out which drives are in which set through the bios, but what is the most likely layout? 5 + 5 + 2 or 5 + 1 + 5 + 1 or something else?

    Is there any risk to trying both of these in retagging? What would happen if we retag incorrectly?

    [On the system, the four virtual disks are formed into one filesystem using md and RAID 0.]

  • UCTpro5

    4 Posts

    632

    0

    Posted October 13th, 2014 04:00

    Just a quick update on this, in case anyone else has a similar issue: Dell support were useless.

    In the end, we removed the drives from the array and put them in a different server and used the free ReclaiMe recovery software to identify the RAID array structure, which was RAID50. We also used R-studio to recover most of the data.

    After replacing the disks, we then re-tagged them in the array (without initializing) and the array was able to build. However, due to the corruption on one or more of the disks, the file system wouldn't mount. We ran e2fsck to attempt to sort out the issues, but it was taking so long, we chose to reformat file system.