I have a cluster that seems to never have any issues. I am running OneFS 7.2.1.4. Every morning I run the following command and all seems to be well nearly every morning:
isi status; isi alert
Well today found out that we had 4 hard disk marked as needing replacement on 4 different nodes within the cluster and just this past week if it wasn't for me catching one of the node lite up yellow on the isi status I would not have caught one of the hard disk smart failing. Comparing the event notifications to another cluster that seems to be working well with event notification everything seems to be configured similarly. Does anyone have any suggestions for where I should look next?
Sorry, I misread. I thought you just weren't getting the notifications, but that you found them when you ran the commands.You could try resetting celog, but that may hinder the investigation through the SR, so you may want to check with the TSE working your request before proceeding.
sjones51
252 Posts
2778
0
Posted August 7th, 2017 13:00
Sorry, I misread. I thought you just weren't getting the notifications, but that you found them when you ran the commands.You could try resetting celog, but that may hinder the investigation through the SR, so you may want to check with the TSE working your request before proceeding.
304312 : OneFS: How to reset the CELOG database and clear all historical events https://support.emc.com/kb/304312.