Earlier today, we had a 'predictive failure' from a PowerEdge 720xd server on one of the disks. About 1-2 hours later server rebooted and came back up in about 5 minutes.
I signed on to the server, but it was extremly slow and database team reported that they were not able to start SQL services running on the server. At this time I have noticed that disk now was in failed a failed state
This went on for about 45 minutes and suddenly server started functioning as normal again.
There is no memory dump, but can see the event id 41 in system logs.
In Dell OME, I have also noticed ID 2048 (failed disk) and after that I see some warning messages, 2346, 2057 and 2123.
I find it be coincidental for server to crash/reboot right after predictivve failure.
Forgot to add, As soon as the server came back up after it crashed, hotspare rebuild has kicked in. To me this looks like the time it took to rebuild the hostspare was the time we were not able to do anything on the server.
Not normal for it to be crippled while rebuilding, but certainly degraded, which can cause some programs not to work very well - or at all. Settings in the controller can be changed to use MORE of the system resources to rebuild - default is 30% - if it has been set to higher than that, it could certainly explain it. Should really never be set to higher than 30%.
Latif_ee7dc4
48 Posts
857
0
Posted December 22nd, 2016 07:00
Forgot to add, As soon as the server came back up after it crashed, hotspare rebuild has kicked in. To me this looks like the time it took to rebuild the hostspare was the time we were not able to do anything on the server.