This post is more than 5 years old
2 Intern
•
143 Posts
1
2866
April 13th, 2015 09:00
Risks associated with initiating a Proactive Hot Spare operation?
I have a customer with a VNX 5300 Block Only system which just experienced a disk failure (reported as "Faulted" in Unisphere and "Disk Removed" in the list of disks). However, based on the Event Logs, it looks like a "Soft Media Error." I noticed that neither of the (2) Hot Spare drives kicked in. I presume this is by design with a Soft Media Error? I do see an option to "Copy to Hot Spare" which I believe kicks off an on-demand Proactive Hot Spare operation. My question is, are there risks associated with kicking off a Proactive Hot Spare operation? The reason I ask is that the VNX Support contract just expired, and we are in the process of renewing it. Because this VNX is in a far-away country, the paperwork is going to take a long time, so I want to minimize risk of data loss caused by this drive failure. Is the EMC best practice to leave the failed disk alone (in the "Soft Media Error" state), or kick off a manual "Copy to Hot Spare" operation? The code level is VNX OE 32 Patch 215 and the disk is part of a RAID 5 storage pool (not RAID group). There are two inactive hot spares of the same type (SAS) available.
Thanks,
Bill


brettesinclair
2 Intern
•
715 Posts
0
April 13th, 2015 15:00
If the disk is faulted, a copy to hotspare operation should have been started automatically, regardless of the cause. (This is assuming the failed disk has a valid candidate to act as a hot spare for it -- size and architecture)
I would start USM and run the "Hardware Replacement" wizard to see if the array is showing the faulted disk as a replacement candidate. It will verify there are no pending transtitions and what state it is in.
If you are some time away from getting support, you may then consider unbinding/deleting one of the hotspares, removing it and using it as a replacement for the faulted disk, providing USM shows it being ok for replacing. This will leave you with one hotspare and a full functional Pool, with no transitions in progress.
dynamox
11 Legend
•
20.4K Posts
•
87.4K Points
0
April 13th, 2015 18:00
HS will not kick in for a drive that is not bound to a pool/RG. "Sparing" is ultimately done at LUN level, individual LUNs are spared to HS. This was more evident when used with traditional RG, not entire 300G drive has to be rebuild if only contained a 10G LUN.
brettesinclair
2 Intern
•
715 Posts
0
April 13th, 2015 19:00
OP said that the faulted disk is part of a Storage Pool. Providing the HS RG is setup, and is the correct size and type, it should be sparing for the faulted disk.
Maybe the Pool was quite full, and thus the disk also. Who knows. Without support or parts, and with an array in a faraway country that hasn't 'protected' itself as it should in a fault condition, I'd be playing it safe.
boyler05
2 Intern
•
143 Posts
0
April 14th, 2015 11:00
Thanks for the great advice. I did not even think about unbinding a hot spare and using USM to "replace" the bad drive. However we are fortunately making good progress in getting the EMC Support contract renewed, so we are going to see if that can be completed in a timely manner, because I'd rather have EMC Support physically on site in the far away country, just in case. I am concerned that the VNX did not automatically kick in one of the hot spares after a soft media error, now that I know it should have. There might be a larger issue. This system had both SP's replaced two years ago (and vault drives rebuilt from scratch by EMC Support) so I am skeptical. Better to have EMC Support. But nonetheless, thanks for the great advice because it may end up being "Plan B".