Unsolved
This post is more than 5 years old
2 Posts
0
150144
November 2nd, 2012 11:00
Bad Disks on MD300
I get the messages below. The part that scares me is Service action (removal) allowed: No What does that mean? I assume that the reason I am getting that on the "Virtual Disk -Hot Spare in Use" notification is because it started using the hot spare. Why Am I getting this service action on the By-Passed Physical Disk? Is it safe to swap out the bad disk? I have a spare.
By-passed Physical Disk
Storage array: SANPNC01
Component reporting problem: Physical Disk in slot 0
Status: Unknown
Location: RAID Controller Module/Expansion enclosure
Component requiring service: 0
Service action (removal) allowed: No
Service action LED on component: No
What Caused the Problem?
The array has detected the presence of a physical disk but is unable to communicate with the physical disk. The Recovery Guru Details area provides specific information you will need as you follow the Recovery Steps.
Caution
Possible loss of data accessibility. Do not remove a component when the Service Action Allowed field in the Details area of this recovery procedure is NO. Removing a component while its Service Action Allowed field is NO may result in temporary loss of access to your data. Refer to the following Important Notes and the Recovery Steps for more detail.
Caution
Electrostatic discharge can damage sensitive components. Always use proper antistatic protection when handling components. Touching components without using a proper ground may damage the equipment.
Important Notes
No data has been lost.
When replacing a physical disk, make sure the replacement physical disk is the same type and has a capacity equal to or greater than the by-passed physical disk.
You can replace the by-passed physical disk while the storage array is performing I/O operations.
The Service Action Allowed field in the Details area indicates whether you can safely remove the component. If the Service Action Allowed field is NO, then the affected component must remain in place until you service another component first.
Recovery Steps
| 1 |
Check the Service Action Allowed status in the Details area. |
|
| If... |
Then... |
|
| Service Action Allowed field is YES |
Go to step 2. |
|
| Service Action Allowed is NO |
No - Stop this procedure and contact your technical support representative. |
|
| 2 |
Obtain a Dell-supported replacement for the by-passed physical disk. |
|
| 3 |
Remove the by-passed physical disk. |
|
| 4 |
Wait 30 seconds and then re-insert the physical disk. |
|
| 5 |
Click the Recheck button to rerun the Recovery Guru. The failure should no longer appear in the Summary area. If the physical disk is still reported as by-passed, contact your technical support representative. |
|
Note:
Additional information on this issue may be available. Please visit the Dell support website at support.dell.com and select your product model. Choose "troubleshooting" as your tool option, then search by this procedure title.
Virtual Disk - Hot Spare in Use
Storage array: SANPNC01
Disk group: 1
Failed physical disk at: enclosure 0, slot 0
Service action (removal) allowed: No
Service action LED on component: No
Replaced by physical disk at: enclosure 0, slot 13
Virtual Disks: SVRPNC01_Data
RAID level: 1
Status: Optimal
What Caused the Problem?
One or more physical disks have failed, and hot spare physical disks have automatically taken over for the failed physical disks. The data on the virtual disks is still accessible. The Recovery Guru Details area provides specific information you will need as you follow the Recovery Steps.
Caution:
Electrostatic discharge can damage sensitive components. Possible loss of data accessibility.
Do not remove a component when the Service Action Allowed field in the Details area of this recovery procedure is NO. Removing a component while its Service Action Allowed field is NO may result in temporary loss of access to your data. Refer to the following Important Notes and the Recovery Steps for more detail.
Caution:
Electrostatic discharge can damage sensitive components.
Always use proper antistatic protection when handling components. Touching components without using a proper ground may damage equipment.
Important Notes
When a hot spare takes over for a failed physical disk, data from the failed physical disk is rebuilt on the hot spare. When you replace the failed physical disk, data is copied back from the hot spare to the new physical disk and the hot spare returns to Standby. You can replace the failed physical disk before rebuild is completed on the hot spare. However, the copyback to the new physical disk will not occur until the rebuild has completed.
From the Summary tab, click the Virtual disks and disk groups link and look at the virtual disk icons for the affected virtual disks. If any virtual disks display an Operation in Progress icon, rebuild is still taking place on the hot spare. If all virtual disk icons are Optimal, rebuild is completed.
Depending on how many hot spares you have created in the storage array, a virtual disk could remain Optimal and still have multiple failed physical disks (each one being covered by a hot spare).
Make sure the replacement physical disks are of the same type and have a capacity equal to or greater than the failed physical disks.
You can replace the failed physical disks while the affected virtual disks are receiving I/O.
The Service Action Allowed field in the Details area indicates whether you can safely remove the component. If the Service Action Allowed field is NO, then the affected component must remain in place until you service another component first.
Recovery Steps
| 1 |
Check the Service Action Allowed status in the Details area for the failed physical disk. |
|
| If... |
Then... |
|
| Service Action Allowed is YES |
Go to step 2. |
|
| Service Action Allowed is NO |
Answer the following question: Are there other problems being reported in the Summary area? Yes - Fix these problems first and then return to this procedure after clicking Recheck. No - Stop this procedure and contact your technical support representative. |
|
| 2 |
If... |
Then... |
| You want to replace the failed physical disk with a new physical disk |
Remove the physical disk (its status LED may be amber flashing). Wait 30 seconds, and then insert the new physical disk. Result: The affected virtual disks change to an Operation in Progress icon (in the Disk groups and virtual disks link) as the rebuild/copyback operations take place. When rebuild/copyback is completed, the virtual disks return to Optimal, and the hot spare physical disk returns to Standby. Repeat above steps for each failed physical disk. Click the Recheck button to rerun the Recovery Guru to ensure that the failure has been fixed. |
|
| You want to utilize an existing unassigned physical disk or the in-use hot spare to replace the failed disk in the disk group |
Click on the Modify tab and then select Replace Physical Disks. Under Failed and Missing Physical Disks, select the physical disk that you would like to replace Under Available replacement physical disks, select the physical disk that you would like to use to replace the failed or missing physical disk Click on Replace Physical Disk. Result: If you choose an unassigned physical disk as a replacement, the affected virtual disks change to an Operation in Progress icon (in the Disk groups and virtual disks link) as the rebuild/copyback operations take place. When rebuild/copyback is completed, the virtual disks return to Optimal, and the hot spare physical disk returns to Standby. If you choose the Hot Spare as a replacement for the failed or missing physical disk, the Hot Spare role will be changed to Assigned. A new Hot Spare would need to be assigned if that functionality is desired Repeat above steps for each failed physical disk Click the Recheck button to rerun the Recovery Guru to ensure that the failure has been fixed. |
|
Note:
Additional information on this issue may be available. Please visit the Dell support website at support.dell.com and select your product model. Choose "troubleshooting" as your tool option, then search by this procedure title.


DELL-Sam L
Community Manager
•
8K Posts
•
357 Points
0
November 2nd, 2012 15:00
Hello UGAmike,
“Service action (removal) allowed: No” what that means is that the drive has most likely failed but it is continues to establish a connection with the MD. What you can do is to manually fail the drive using SMcli but it may fail if the array already see the drive as failed .
To manually set a physical disk to offline status you have to use SMcli, the command line interface for the MD3xxx series storage arrays.
smcli -n enclosureName -p enclosurePassword -c "set physicalDisk [enclosureID,slotID] operationalState=failed;"
EXAMPLE:
You have an MD3200i named "DellArray" with an attached MD1220 with the enclosure ID of 1. Slot 7 in the MD1220 is showing an impending physical disk failure. Your password for the enclosure is "Passw0rd" The command to manually fail physical disk 1,7 would look like this:
Smcli -n DellArray -p Passw0rd -c "set physicalDisk [1,7] operationalState=failed;"
If you do not have a password set on the array, leave out the password segment
Smcli -n DellArray -c "set physicalDisk [1,7] operationalState=failed;"
Let us know if have any other questions.