This post is more than 5 years old
4 Posts
0
119684
August 5th, 2014 03:00
Poweredge 2900 showing Predicted Failure after replacing drive
Hello,
I have posted this same message under the Disk Drives forum, then I realized there is a specific forum for Server HDDs so I am copying it here.
We have a Poweredge 2900 with 2x500Gb HDs on RAID-1. One of the disks failed and we have replaced it with an identical disk as the one working fine (same brand, Seagate model number, etc.), though without a Dell PN.
After replacing it, the controller has started to rebuild the RAID but it shows "Predicted Failure=Yes" for this new drive. The disk is currently still under rebuild.
One thing to notice is that apparently the Perc firmware is outdated. OpenManage shows that the current Perc 6/i firmware version is 6.0.2-0002, and the minimum required firmware version is 6.2.0-0012. If I go to the Dell support site and enter the server's service tag it shows me an available download for the Perc firmware with version 6.3.3-0002.
So the questions are:
- Is it usually ok to use non-Dell certified drives if they are the same brand and model as the other disk?
- How could I determine the source for the Predicted Failure warning?
- Could this have something to do with the outdated Perc firmware? Are there any risks associated with upgrading the Perc firmware?
Many thanks in advance for your help!
--
Julio.


DELL-Bo Pham
2 Intern
•
261 Posts
0
August 22nd, 2014 14:00
Updating thread for our forum followers. Original poster found that the replacement drive used was indeed bad. He is replacing it later with a good drive and will proceed with updating the server accordingly.
DELL-Bo Pham
2 Intern
•
261 Posts
0
August 5th, 2014 10:00
Hi Julio,
Yes, the firmware update is highly recommended along with latest driver. It is also important to update the BIOS and BMC to allow to proper communication with the updated PERC FW. The 6.3.3. FW is correct. The reason why your OMSA application is indicating 6.2.0 is that there is a library within your old OMSA installation which referenced the latest 6.2.0 that was current during that version build. It is only comparing what is on the PERC with what it knew at the time. Updating to the latest OMSA version will resolve the discrepancy.
With any FW update there can always be a possible hardware issue during the flash, but it is rare. Always ensure you have a good backup prior to any system update. Do not update FW while the system if rebuilding the array. Let it finish first.
As for the predictive failed state, please ensure that you took the affected drive offline within the OMSA tool or from the RAID controller BIOS before replacing the drive. There are other possible issues. The non-certified drive will not have the proper Dell Firmware to communicate with PERC controller and may cause further issues. We recommend that you use Dell certified drives which are available through Dell or many third party Vendors. Lastly, the replacement drive could be defective or your RAID may be corrupted.
Depending on the controller type, if it has the ability to export the controller log, then you may send it to me for review. I’ll be glad to take a look at it for you.
Please let me know if you need any links to the items referenced above.
julio.gutierrez
4 Posts
0
August 5th, 2014 11:00
Hello Bo,
Many thanks for your quick response and your offer to check the log!
The old drive was not taken offline prior to inserting the new; but actually I could not have done it as OMSA didn't show it at all. If I inserted the failed drive, the server would try to rebuild for 5-10 minutes, then the drive would show a "taking disk off" state for some seconds (do not know the specific state in English as my OMSA is in Spanish) and then the physical disk would not be available anywhere.
Anyway, the new drive has completed the rebuild and it still shows Predicted Failure=Yes. I have exported the controller log; please let me know how to send it to you as I don't seem to find a tool within the forum to attach files to this message. I see a couple of things that may be relevant (being the first time I have a look at logs like this I could be all wrong though):
This is what it shows when the new disk was inserted:
08/05/14 9:52:50: PD Flags State Type Size S NCQ Vendor Product Rev P C ID SAS Addr Port Phy DH
08/05/14 9:52:50: --- -------- ----- ---- -------- - --- -------- ---------------- ---- - - -- ---------------- ---- --- --
08/05/14 9:52:50: 0 f0400005 00020 00 3a38602f 1 0 ATA GB0500C4413 HPG1 0 0 00 1221000000000000 00 00 0b
08/05/14 9:52:50: 1 f0400005 00020 00 3a38602f 1 0 ATA ST3500630NS 3BKS 0 0 01 1221000001000000 01 01 0c
08/05/14 9:52:50: 20 00400005 00020 0d 00000000 0 0 DP BACKPLANE 1.05 0 0 20 5001e0f0393fe600 09 08 09
08/05/14 9:52:50: 100 00400005 00020 03 00000000 0 0 LSI SMP/SGPIO/SEP 0396 0 0 ff 0 00 ff 00
08/05/14 9:52:50: PhyId 0 Sas 1221000000000000 Type 1 IsSata 1, Smp 0:0
08/05/14 9:52:50: PhyId 0 Sas 1221000001000000 Type 1 IsSata 1, Smp 0:0
The disk with serial number STxxxxx is the existing DELL one; the disk with GBxxxxx is the new disk we just inserted today.
Next, drive is identified as non-certified, rebuild starts but I see a couple of potential error messages (last 4 lines).
08/05/14 9:52:50: EVT#14534-08/05/14 9:52:50: 247=Inserted: PD 00(e0x20/s0) Info: enclPd=20, scsiType=0, portMap=00, sasAddr=1221000000000000,0000000000000000
08/05/14 9:52:50: EVT#14535-08/05/14 9:52:50: 236=PD 00(e0x20/s0) is not a certified drive
08/05/14 9:52:50: EVT#14536-08/05/14 9:52:50: 114=State change on PD 00(e0x20/s0) from UNCONFIGURED_BAD(1) to UNCONFIGURED_GOOD(0)
08/05/14 9:52:50: EVT#14537-08/05/14 9:52:50: 114=State change on PD 00(e0x20/s0) from UNCONFIGURED_GOOD(0) to OFFLINE(10)
08/05/14 9:52:50: EVT#14538-08/05/14 9:52:50: 106=Rebuild automatically started on PD 00(e0x20/s0)
08/05/14 9:52:50: EVT#14539-08/05/14 9:52:50: 114=State change on PD 00(e0x20/s0) from OFFLINE(10) to REBUILD(14)
08/05/14 9:53:05: process_raid_msg: Unknown DCDB Opcode 85
08/05/14 9:53:05: Rdm a084a600 Cdb 85 09 0e 00 00 00 01 00 99 00 00 00 00 00 2f 00
08/05/14 9:53:05: process_raid_msg: Unknown DCDB Opcode 85
08/05/14 9:53:05: Rdm a084ec00 Cdb 85 09 0e 00 00 00 01 00 99 00 00 00 00 00 2f 00s)
I don't know whether this will help; otherwise let me know how can I send the complete log file to you.
Regarding the FW updates: Could you please post links to where I can download the proper BIOS and BMC files? Also, is there any specific order in which the updates need to be applied? (e.g. BIOS first, then BMC, then Perc).
Also, where could I download the latest OMSA version?
Again, many thanks for your help. Really appreciated.
--
Julio.
DELL-Bo Pham
2 Intern
•
261 Posts
0
August 5th, 2014 15:00
Glad I am able to help Julio. Those errors are communication related which may or may not be a problem. I sent you an email with the update links for your system. You can attach the LSI log and send it back to me so I can view the log in full. For now, you may try to reboot the server, or restart the 5 OMSA services to refresh the communication of the drives. There are 4 OMSA services that start with “DSM SA …” and one that is called “mr2kserv. Close your OMSA browser and stop them all those sevices, then restart them one at time in the same order. Then check your drive status again.
You may apply the updates as follows:
Please let us know if you have any questions.
julio.gutierrez
4 Posts
0
September 1st, 2014 03:00
We received a Dell-certified drive and inserted it in the disk bay. After rebuild and 2-day operation, OMSA still shows all ok.
Issue has been resolved.
My most sincere thanks to Bo. This could not have been resolved without your help!
--
Julio.