Unsolved
This post is more than 5 years old
8 Posts
0
117609
September 14th, 2015 12:00
RAID 5 not automatically rebuilding
Dell PowerEdge T610, with eight 300GB drives
PERC 6/i, RAID 5, one of the drives is a hot spare
------------------------------------------------------------------
In the past, when a drive has failed, the server keeps running fine thanks to the hot spare. I order a replacement drive, insert it, and the rebuild starts automatically. No other action required, easy peasy.
This time is different. A couple days ago drive 1 failed (blinking amber light). The server continues to run just fine. As per my routine I ordered a new replacement drive with the same exact specs. I insert it, and after a minute or two, the amber light starts flashing. I re-seated it to make sure it was connecting properly, but again after a minute or two the amber light starts to flash.
In OpenManage, under Storage\Perc6\Connector 0\Enclosure\Physical Disks, the drive is not shown. The list shows:
Physical Disk 0:0:0
Physical Disk 0:0:2
Physical Disk 0:0:3
However just below that Physical Disks folder there is a new "Physical Disks" sub folder that contains one item: Physical Disk 0:33 , I believe that is the replacement drive that is sitting in limbo (?) because the serial number matches. Also, when I go to "View Slot Occupancy Report", I see 9 slots total... slot 1 shows as being empty (this is where the new drive is physically located) , and there's a "slot 33" at the far right position.
Apparently I have two action options with this drive: Assign as Global Hot Spare, or, Clear.
What should I do? This is all new to me, like I said the rebuild was always automatic before. I don't know why it's acting different this time. All the other drives seem healthy.
Any ideas? I am extremely grateful for your help. The server is running fine atm, but of course I can't afford to lose another drive here.
Thanks,
-Dave


david.berry
8 Posts
0
September 14th, 2015 13:00
Thanks for the reply. I'll execute Assign as Global Hot Spare.
On the serial number, I should have been more clear, I meant that the serial number displayed in OpenManage matches the new drive's serial number, thus confirming the identification.
Thanks,
-Dave
theflash1932
11 Legend
•
16.3K Posts
0
September 14th, 2015 13:00
There are many things that can prevent a disk from rebuilding automatically. The option you want is to assign it as a hot-spare to start the rebuild.
(What do you mean it has the same "serial number"? No two drives should have the same serial number.)
theflash1932
11 Legend
•
16.3K Posts
0
September 14th, 2015 13:00
Ah, ok ... that isn't the actual/full SN. No biggie :)
david.berry
8 Posts
0
September 15th, 2015 05:00
Looks like I could use some more help on this. Here's the current status (I probably should have shared these screen shots yesterday)... notice how the disk is not included with the other drives (it's supposed to be Disk 0:0:1)...
edit: looks my screen shots didn't work. but anyway, in OMSA, under Storage, Perc6, Connector 0, Enclosure, there are two 'Physical Disks' folders (should only be one). The first one contains disks 0:0:0, 0:0:2, 0:0:3 (notice disk 0:0:1 is missing, that's the one that failed). In the other Physical Disks folder, there is this Disk 0:33 which is the replacement disk. Its state is READY. I assigned it as a Global Hot Spare thinking that would put it where it belongs but no luck. Its only available task atm is 'Unassign Global Hot Spare".
Oh, and its amber light is still blinking :(
As always I am grateful for any/all assistance! Thanks!
-Dave
david.berry
8 Posts
0
September 17th, 2015 10:00
If anyone is reading this, I'd be grateful for any help. I still can't figure it out.
Thanks,
-Dave
david.berry
8 Posts
0
September 17th, 2015 11:00
david.berry
8 Posts
0
September 17th, 2015 11:00
david.berry
8 Posts
0
September 17th, 2015 11:00
Thanks for your reply, and tip on the screen shots.
In the first one, you can see 3 disks (supposed to be 4). One is missing (0:0:1) which is the failed drive. To recap, I ordered a brand new replacement with identical specs, inserted it into slot 1, and after a minute or two its amber light started to blink. I thought it was a bum disk, so I ordered yet another replacement, but I'm still seeing the same behavior. The RAID 5 array is not automatically rebuilding unlike prior situations.
In the second screenshot, you can see the replacement disk is sitting by itself, in another 'Physical Disks' folder, identified as disk 0:33 (?)
Any help is much much appreciated. Thx.
theflash1932
11 Legend
•
16.3K Posts
0
September 17th, 2015 11:00
I think the screenshots would be helpful. You can't paste images. Click on the Insert/Edit Media icon in the middle of the bottom row of the editor window and upload/attach a saved image file.
theflash1932
11 Legend
•
16.3K Posts
0
September 17th, 2015 20:00
Honestly, I've seen the disk 0:33 before (not on my servers), but because everything seems to work normally otherwise, I've never done a lot of troubleshooting with anyone.
As far as the rebuild not being automatic, that is perfectly normal. Assign it as a hot-spare, and as long as it rebuilds and everything shows healthy, the 0:33 thing doesn't seem to be a big deal. I would suggest you update the system firmware (iDRAC, BIOS, PERC, backplane, etc.) to possibly address the disk ID issue. You could post a controller log to see if the log can shed any light as to why it is getting such a strange ID.
david.berry
8 Posts
0
September 18th, 2015 06:00
I appreciate your help. I guess I'm still having trouble understanding this. I've had disk failures in the past on this server, and all I had to do was swap out the bad drive with a replacement and the rebuild would just start by itself. Literally no further action required, no OMSA, nothing.
Some more detail in case it helps:
Prior to this issue, there was no disk 0:33. In OMSA, disk 1 appeared in the list as 0:0:1 along with the other 3 drives. When disk 1 failed, the LCD display showed "E1810 hard drive 1 fault". In OMSA, the disk was gone, as if it did not exist. The disk itself had a flashing amber status light. The activity led was not lit. I believe this all points to the disk indeed failing?
So as you know I replaced the drive... When I insert the new drive, its green activity LED blinks/flashes for ~15 seconds, then stays on solid. After about a minute or so, its status LED starts amber flashing, the activity LED remains green solid, and the LCD displays "E1810 hard drive fault". However in OMSA, the disk 0:33 appears, healthy, but separate from the other 3 disks.
So as you can see, it seems to be conflicting information. For now, I have assigned 0:33 as a Global Hot Spare, but I'm not convinced that is actually a 'member' of that group of disks, since it is listed separately in OMSA. And of course the "hard drive fault" indication is a huge concern... I have tried 2 replacement disks, both brand new, and both have behaved the same.
I hope this additional detail might help? I apologize if I'm belaboring it, but I am very concerned because this is a our primary production server. Again, I am very grateful for any/all help.
Thanks, -Dave
DraganK
1 Rookie
•
36 Posts
0
October 13th, 2015 03:00
Hello,
I had today exactly the same issue on my PE2950 server with PERC 6i.
Disk 0:0:1 failed and replaced with new one promoted to HS that is shown as 0:33. Rebuild finished but the new one still blinking amber. Message backplane degraded appeared in OMSA.
There are some notes about log clearing using DSET and I did it.
DSET log clear didn't solve the issue, amber still blink and backplane degraded still exists.
I shutdown server unplug AC connector for a minute and plug it again, start the server and everything was just fine. Disk shown as 0:33 became 0:0:1 again, according to physical position.
Conclusion:
This worked for me
Good luck