This post is more than 5 years old

4 Posts

20530

November 30th, 2014 01:00

How to recover from reported failure of multiple drives while system powered off for a couple of weeks

Hello,

I had a logical drive failure on a PowerEdge 2600 I recently inherited and I am hoping to get some suggestions on how to fix it.

System was configured with 6 disk drives as a single RAID 5 logical drive with a spare. It was running Windows Server just fine. I was configuring the system and installing software in my spare time so I was not leaving it running constantly.

After it had been down for a couple of weeks, I powered it back up and something had gone wrong with my logical drive:

   HA -0 (Bus 8 Dev 8) PERC 4/Di Standard FW 252D DRAM-128MB
   Battery module is present on adapter
   1 Logical Drives found on the host adapter.
   1 Logical Drive(s) Failed
   1 Logical Drive(s) handled by BIOS
   Configuration of NVRAM and drives mismatch(Normal mistmatch)
   User Configuration....

All of the drives seem to spin up and they all had green LEDs lit.

However, according to the PERC Configuration Utility Configure > View/Add to DISK Configuration tool there were problems with multiple drives:

   ID 0 ONLIN A00-00
   ID 1 FAIL A00-01
   ID 2 READY
   ID 3 ONLIN A00-03
   ID 4 FAIL A00-04
   ID 5 FAIL A00-02
   ID 6 PROC

And this is what the PERC Configuration Utility Configure > View/Add to NVRAM Configuration tool showed:

   ID 0 FAIL A00-00
   ID 1 FAIL A00-01
   ID 2 READY
   ID 3 FAIL A00-03
   ID 4 FAIL A00-04
   ID 5 FAIL A00-02
   ID 6 PROC

I went searching through the forum archives and found a post that helped me to address the disk/NVRAM mismatch issue.

Unfortunately, now that is fixed but I get amber blinking LEDs on drives 1, 4, and 5 and still get the dreaded "1 Logical Drive(s) Failed" message.

It seems really suspect to me that THREE drives would fail during the same brief interval the machine was off.

I am a total newb with this system and have no idea how to resolve the problem and resurrect my system. Can anyone please give me any pointers?

Thanks so much,
Jon

12 Elder

 • 

6.2K Posts

November 30th, 2014 15:00

Hello Jon

I went searching through the forum archives and found a post that helped me to address the disk/NVRAM mismatch issue.

What did you do? You likely overwrote one of the above configurations that you posted. The controller stores the array configuration and drive status locally in NVRAM. The information is also stored on each of the drives. If the information on the drives does not match the controller then you get the mismatch error. If the information on the drives do not match each other then the drives that are out of sync are put into a failed state.

Did you choose to save the configuration from disk?

Your only option besides a data recovery company is a retag. You should write down or take screen shots of all of the array information before you do anything else. A retag is when you delete the current array and then recreate it with all of the same drives and settings. It will recreate the metadata tags that describe the array. Be sure to not initialize the array or it will wipe the data.

There are a lot of posts about retagging, what it is, and how to do it. If you want more information on it then please search the forums.

Thanks

4 Posts

November 30th, 2014 19:00

Truthfully, I am not sure exactly what I did. I think I saved the configuration from disk. There is no longer a mismatch error and the current status is thus (which is what the disk configuration was previously):

   ID 0 ONLIN A00-00
   ID 1 FAIL A00-01
   ID 2 READY
   ID 3 ONLIN A00-03
   ID 4 FAIL A00-04
   ID 5 FAIL A00-02
   ID 6 PROC

I will research retag and give that a whirl; thank you for the pointer.

Any idea how the drives end up in that out of sync state? The system was always shut down cleanly and never was endured anything like a power failure. Without knowing why it failed, I am afraid that even if I resurrect that raid set the same thing might happen again.

Thank you,
Jon
 

4 Posts

December 10th, 2014 22:00

I rettagged as suggested which resurrected the system.

When I went to recreate the logical drive, PERC said drive 2 had a fault. Drive 5 was hot spare so I swapped the two drives and then recreated the logical volume.

Only remaining cleanup is to replace the failed drive so I have a hot spare.

Thanks so much to Daniel Mysinger for the pointer!

8 Posts

December 19th, 2014 03:00

Here's how you can recover back all your data from pen drive, internal hard drive, external hard drives and so on - 

http://youtu.be/XNsyNkc_CTk

6 Operator

 • 

1.8K Posts

December 19th, 2014 11:00

"Any idea how the drives end up in that out of sync state?"

Aside from firmware issues

By probability.....

Drive/raid adapter electronic component out of spec, intermittent component failure 

Drive physical error

Power anomaly getting past power filter/UPS/power supply.

EMF , industrial machines, lightning storms.

Someone screwing around with the server without your knowledge

Particle from space

Bad vibes at the office

The Iranians getting back at us for Stuxnet

No Events found!

Top