I have a disk in a raid 10 virtual disk that is failing. The physical disk caddy has an amber light on it and in the perc display it shows with an orange yield sign with an exclamation mark in it. I assume that means it is failing even though it is still running and not yet degrading the virtual disk. I have a new spare now to replace it. I don't have a spare slot in the chassis. If I did I would insert the new disk, make it a hot spare for the virtual disk, then pull the failing disk. Since I don't have a spare slot. what is the best way to approach this. In other words, do which series of commands in which order? Do I just take the failing disk off line, pull it, replace with new, mark it as hot spare? I'd kind of like assurance that I'm doing it in the right order before going ahead
The disk is a member of Controller0, Virtual disk1. It is disk#6 on enclosure#32
And indeed queerying the disk gives me the same serial number as the one that shows in idrac with an amber exclamation mark
[root@CCV-VMware-Lam:/opt/lsi/perccli] ./perccli /c0/v1 show all CLI Version = 007.0529.0000.0000 Sep 18, 2018 Operating system = VMkernel 6.7.0 Controller = 0 Status = Success Description = None
/c0/v1 : ======
---------------------------------------------------------------- DG/VD TYPE State Access Consist Cache Cac sCC Size Name ---------------------------------------------------------------- 1/1 RAID10 Optl RW Yes RWBD - OFF 1.635 TB SAS15K ----------------------------------------------------------------
------------------------------------------------------------------------------ EID:Slt DID State DG Size Intf Med SED PI SeSz Model Sp Type ------------------------------------------------------------------------------ 32:10 10 Onln 1 558.375 GB SAS HDD N N 512B ST600MP0005 U - 32:11 11 Onln 1 558.375 GB SAS HDD N N 512B ST600MP0005 U - 32:6 6 Onln 1 558.375 GB SAS HDD N N 512B ST600MP0005 U - 32:7 7 Onln 1 558.375 GB SAS HDD N N 512B ST600MP0005 U - 32:8 8 Onln 1 558.375 GB SAS HDD N N 512B ST600MP0005 U - 32:9 9 Onln 1 558.375 GB SAS HDD N N 512B ST600MP0005 U - ------------------------------------------------------------------------------
VD1 Properties : ============== Strip Size = 64 KB Number of Blocks = 3512991744 VD has Emulated PD = No Span Depth = N/A Number of Drives Per Span = N/A Write Cache(initial setting) = WriteBack Disk Cache Policy = Disk's Default Encryption = None Data Protection = None Active Operations = None Exposed to OS = Yes Creation Date = 07-03-2016 Creation Time = 01:18:54 PM Emulation type = default Cachebypass Mode = Cachebypass Disable Is LD Ready for OS Requests = Yes SCSI NAA Id = 61418770601b16001e703c3e0950c5f9 SCSI Unmap = No
[root@CCV-VMware-Lam:/opt/lsi/perccli] ./perccli /c0/e32/s6 show all CLI Version = 007.0529.0000.0000 Sep 18, 2018 Operating system = VMkernel 6.7.0 Controller = 0 Status = Success Description = Show Drive Information Succeeded.
Drive /c0/e32/s6 : ================
------------------------------------------------------------------------------ EID:Slt DID State DG Size Intf Med SED PI SeSz Model Sp Type ------------------------------------------------------------------------------ 32:6 6 Onln 1 558.375 GB SAS HDD N N 512B ST600MP0005 U - ------------------------------------------------------------------------------
Drive /c0/e32/s6 - Detailed Information : =======================================
Drive /c0/e32/s6 State : ====================== Shield Counter = 0 Media Error Count = 0 Other Error Count = 0 Drive Temperature = 37C (98.60 F) Predictive Failure Count = 1 S.M.A.R.T alert flagged by drive = Yes
Drive /c0/e32/s6 Device attributes : ================================== SN = S7M0QBXX Manufacturer Id = SEAGATE Model Number = ST600MP0005 NAND Vendor = NA WWN = 5000C5008F5CE770 Firmware Revision = VT31 Firmware Release Number = N/A Raw size = 558.911 GB [0x45dd2fb0 Sectors] Coerced size = 558.375 GB [0x45cc0000 Sectors] Non Coerced size = 558.411 GB [0x45cd2fb0 Sectors] Device Speed = 12.0Gb/s Link Speed = 12.0Gb/s Write Cache = Disabled Logical Sector Size = 512B Physical Sector Size = 512B Connector Name = 00
Drive /c0/e32/s6 Policies/Settings : ================================== Drive position = DriveGroup:1 Enclosure position = 1 Connected Port Number = 0(path0) Sequence Number = 2 Commissioned Spare = No Emergency Spare = No Last Predictive Failure Event Sequence Number = 6513 Successful diagnostics completion on = N/A SED Capable = No SED Enabled = No Secured = No Cryptographic Erase Capable = No Locked = No Needs EKM Attention = No PI Eligible = No Certified = Yes Wide Port Capable = No
Port Information : ================
----------------------------------------- Port Status Linkspeed SAS address ----------------------------------------- 0 Active 12.0Gb/s 0x5000c5008f5ce771 1 Active 12.0Gb/s 0x0 -----------------------------------------
With the drive showing as an online Predicted Failure, you will need to force that drive offline (page 25 here). prior to removal. After that you should be able to remove the problem drive, wait about 20 seconds and insert the replacement. It should start the rebuild automatically, if not then you can assign it as a hotspare with commands found on page 34 here.
Let me know how it goes.
DELL-Chris H
Social Media and Communities Professional
Dell Technologies | Enterprise Support Services
#IWork4Dell
Did I answer your query? Please click on ‘Mark as Accepted Answer’. ‘Thumbs up’ the posts you like!
Hello, I am aware that this post is particularly old but I have exactly the same problem on an H840 card and 15 disks in RAID 5 One disk has problems : Drive has flagged a S.M.A.R.T alert : yes Adapter 1 Enclosure Device ID: 251 Slot Number: 5
Is it possible to do the same thing with "megacli" that I usually use ? megacli -PDOffline -PhysDrv[251:5] -a1 Then put the disk OffLine, then remove it mechanically, then put the new one in its place. The rebuild starts by itself? I see on the net that you have to put the disk in "missing state" then in "removable state" before removing it, is it really useful in the case of a replacement? Thanks for your help
In regards to a Predicted Failure drive, which is in an ONLINE state, you indeed need to OFFLINE that drive prior to replacing it. This is to ensure that the bad blocks on the Predicted Failure drive don't get moved to the replacement drive. So you would offline it, replace it, then if the rebuild doesn't automatically start you can set it as a Hotspare to start it.
Let me know if this answers your question.
DELL-Chris H
Social Media and Communities Professional
Dell Technologies | Enterprise Support Services
#IWork4Dell
Did I answer your query? Please click on ‘Mark as Accepted Answer’. ‘Thumbs up’ the posts you like!
My question was more to know if in this case (the same as the one mentioned at the beginning of this post : replacement of a defective disk), it was enough to only pass it in Offline* before removing it and before putting the new one.
* With this megacli command : megacli -PDOffline -PhysDrv[251:5] -a1
That looks correct to me. Also, just on a side note, if you ever need to offline a drive, when you don't have access to it via cli or gui, you can power off the server, then remove the drive, then power up the server, then add a replacement drive when back in the OS.
DELL-Chris H
Social Media and Communities Professional
Dell Technologies | Enterprise Support Services
#IWork4Dell
Did I answer your query? Please click on ‘Mark as Accepted Answer’. ‘Thumbs up’ the posts you like!
DELL-Chris H
7 Practitioner
•
9682 Posts
•
48046 Points
9834
1
Posted December 4th, 2019 06:00
Billeuze,
With the drive showing as an online Predicted Failure, you will need to force that drive offline (page 25 here). prior to removal. After that you should be able to remove the problem drive, wait about 20 seconds and insert the replacement. It should start the rebuild automatically, if not then you can assign it as a hotspare with commands found on page 34 here.
Let me know how it goes.
DELL-Chris H
Social Media and Communities Professional
Dell Technologies | Enterprise Support Services
#IWork4Dell
Did I answer your query? Please click on ‘Mark as Accepted Answer’. ‘Thumbs up’ the posts you like!