I have several Dell R730 the are "losing" hard disks. I mean that, the os says "disk reading error" the controller put the disks in failed state. If I reboot the controller says I have foreign data and I import it.
But if I extract the disks and I do a smart test I see the disks are perfectly working.
So it seems the controller is getting crazy or the drive cage.
So the question is: which part do I need to repair?
If you have downtime for the server that you can update iDRAC, BIOS, and PERC Then, you can look at the SEL log to see if there are drive-related errors or if there are warnings such as predictive failure. Then, if you have a backup, you can try replacing it to see if the problem is with the hardware. If you have known good parts that you can try to understand which part is the issue by doing X tests, that is, by replacing the parts with known good parts.
Hope that helps!
DELL-Erman O
Social Media and Communities Professional
Dell Technologies | Enterprise Support Services
#IWork4Dell
Did I answer your query? Please click on ‘Mark as Accepted Answer’. ‘Thumbs up’ the posts you like!
I think you are suspecting HDD carrier. This is also a possibility if the Drives do not fit the backplane properly. I would check by doing onsite troubleshooting. You can check the latches.
DELL-Erman O
Social Media and Communities Professional
Dell Technologies | Enterprise Support Services
#IWork4Dell
Did I answer your query? Please click on ‘Mark as Accepted Answer’. ‘Thumbs up’ the posts you like!
Apologies for the interjection, I am having the same issue with an r710. It has been driving me batty for about a month now. Just the last 3.5" slot. Drive failed. 2 minutes later, operating normally. Rinse and repeat every 10 minutes. Ubad, Ugood, online, rebuild, consistency check, all complete. Get home the next day, Yellow lcd, yelllow blink, same. Another oddity, each time it "fails" out of the array and appears back as foriegn, it flip flops from being seen as enclosure 0 slot 5 to just slot 5, no enclosure. Smart self test, short test and extended test passed repeatedly. Full test in the support assist bootable diagnostics passes. No errors in syslog. Taken drive out, no interface on these carriers, backplane out, cleaned everything, checked cable routes, power supplies, all to no avail. Only idea I have left is to break the array and use the manufacturers stand alone diag tools on another system, but that takes more than a few days on a 10 TB. Ideas? (And yes, all firmware has been updated, though I have never seen one for this backplane).
In your case, @alksj460, if diagnostic test out of the OS passes, then probably there are some logical issue on the array.
Which controller is installed? You can verify if there is a sort of check consistency in the BIOS of the controller or via Open Manage Server Administrator.
You said that all firmware are updated, did you check also hard drive firmware?
Thanks
Marco
DELL- Marco B
Social Media and Communities Professional
Dell Technologies | Enterprise Support Services
#IWork4Dell
Did I answer your query? Please click on ‘Mark as Accepted Answer’. ‘Thumbs up’ the posts you like!
It has a PERC H700 with 512MB cache, battery backed. I have run the consistency checks and the rebuilds thru both the boot interrupting bios interface, and via the perc cli utility. As to the drive firmware, no, there have been no updates put out from the vendor. The drive in question is one of four in a raid 5, two being matched part numbers, and the fourth being a newer model RMA replacement this past spring, at least a few months prior to the issue.Only slot 5 has the issue. I have NOT tried swapping two in the array around, because I do not remember if it can rebuild from that without data loss. (This is a home media and whole home backup, so offloading 26 TB is going to take quite a while, and required an r730xd purchase.) And no data loss is the kicker! Even with this drive supposedly failing numerous times; sometimes filing the SEL in an hour or two; when I set it to good and reinsert it to the VD, It's consistent, patrol reads pass, Smartctl under linux passes all tests, Dell diagnostics pass, and the data is all there, no read or write degradation. No errors in the syslog, SEL just says it failed, then operating normally. I have yet to get OMSA to run on this, but the iDRAC is not logging anything more descriptive. the PERC log has a tad bit more info, but is also like trying to decypher a phone number out of the history of the universe. It is logging SENSE errors, Though I believe these were during boot, mainly b/47/03 and 6/26/00. The b/47/03 looks to be either the cause of or in response to an internal device reset, though as to the requestor, I am unsure. Log snippet below.
It can be caused by the firmware of the control not being up to date. As far as I understand and I see, there is no failed drive or predictive failure drive. So it would be good to check the firmware first. If multiple VDs were created from the PERC BIOS interface and there were drives removed, sometimes PERC can get confused and send an invalid command because it cannot access it. Sense codes warnings may also occur when the drives used are non-certified. I would also like to share this wiki page for sense codes that I think will be useful https://dell.to/3qk2PhK FW update of the hard drive may also be useful. You will see the same warning when you look through OMSA. You can also clear these SEL logs via iDRAC or OMSA. Because SEL logs that are not cleared can be listed to remind you again. I will actually recommend iDRAC, BIOS and PERC controller updates when it can get downtime. You can also check into Drive FW. You can access FWs from here by entering ST https://dell.to/3mxiKYJ
If the you see same unexpected sense warning then please reseat the hard drive and PERC cables.
Hope that helps!
DELL-Erman O
Social Media and Communities Professional
Dell Technologies | Enterprise Support Services
#IWork4Dell
Did I answer your query? Please click on ‘Mark as Accepted Answer’. ‘Thumbs up’ the posts you like!
mgiammarco
1 Rookie
•
29 Posts
1810
1
Posted April 9th, 2022 01:00
Changing hardware was not useful. Doing all updates was not useful.
So I went to the bios and changed random options:
- enabled logical processor idling
- disabled sdcard readers
Now the server is perfectly stable and it is a passed a month.
I have other server with sdcard enable and they do not lose disks.
Perhaps saving the bios had cleaned some data.
Mario