Unsolved
46 Posts
0
1834
December 31st, 2020 02:00
X200 both boot drives wear_life threshold exceeded
Hi,
We are using a cluster of IQ6000X and X200 nodes.
Recently we had a power failure on site, our 30kVA UPS was holding but it went offline just before the power came back.
After powering on the cluster node no 8 (X200) reported 2 critical errors:
Drive at Internal J4/ad7 wear_life threshold exceeded: 100 (Threshold: 95). Please schedule drive replacement immediately.
Drive at Internal J3/ad4 wear_life threshold exceeded: 100 (Threshold: 95). Please schedule drive replacement immediately
I found documentation how to replace the boot drive but when I ran the commands to detect witch drives are failing I got that both drives are working:
gyar-8# atacontrol list
ATA channel 0:
Master: no device present
Slave: no device present
ATA channel 1:
Master: no device present
Slave: no device present
ATA channel 2:
Master: ad4 Serial ATA v1.0 II
Slave: no device present
ATA channel 3:
Master: no device present
Slave: ad7 Serial ATA v1.0 II
ATA channel 4:
Master: no device present
Slave: no device present
ATA channel 5:
Master: no device present
Slave: no device present
gyar-8# gmirror status
Name Status Components
mirror/root1 COMPLETE ad7p5
ad4p5
mirror/keystore COMPLETE ad7p12
ad4p11
mirror/mfg COMPLETE ad7p9
ad4p9
mirror/journal-backup COMPLETE ad7p8
ad4p8
mirror/var1 COMPLETE ad7p7
ad4p7
mirror/var0 COMPLETE ad7p6
ad4p6
mirror/root0 COMPLETE ad7p4
ad4p4
mirror/var-crash COMPLETE ad4p10
Could someone tell me if I really need to replace the drives? Do I need to replace both? What would be the procedure in this case when both drives are failing? (in the document there is no description for this case).
We do not have support anymore, so it's on me to solve this issue.
Many thx


DELL-Josh Cr
Community Manager
•
9.7K Posts
•
43.3K Points
0
December 31st, 2020 08:00
Hi Andrew,
This thread has some good information about replacing the drives. https://dell.to/3rLAMr9
AndrewF76
46 Posts
0
January 1st, 2021 01:00
Hi Josh,
Thx for the quick response.
I have read that thread, but I couldn't find info about replacing both boot drives.
From that threat I ran these 2 commands too and got these responses:
Isilon OneFS v7.1.1.11
gyar-8# isi_for_array -s isi_radish -a /dev/ad* | grep -E "Percent Life" | grep -v Used | awk '{print $1 $2 $3 $4 sprintf( "%d","0x"$9 )}'
gyar-5:PercentLifeRemaining:13
gyar-5:PercentLifeRemaining:16
gyar-7:PercentLifetimeLeft:77
gyar-7:PercentLifetimeLeft:77
gyar-8:PercentLifetimeLeft:1931266
gyar-8:PercentLifetimeLeft:1934950
gyar-11:PercentLifetimeLeft:81
gyar-11:PercentLifetimeLeft:81
gyar-12:PercentLifetimeLeft:79
gyar-12:PercentLifetimeLeft:78
gyar-8# isi_radish -q /dev/ad[0,1,2,3,4,7]
Internal J3/ad4 is NETLIST SSD 8GB-001 FW:SBFM01.1 SN:5P1171218003000038, 15649200 blks
Internal J4/ad7 is NETLIST SSD 8GB-001 FW:SBFM01.1 SN:5P1171218003000044, 15649200 blks
gyar-8#
Node 8 has reported the failing boot drives, so what exactly means the 1931266 and 1934950 values?
Do I actually need to replace them?
Thx
Andrew
AndrewF76
46 Posts
0
January 4th, 2021 10:00
Hi Josh,
Thank you for your help.
I will try to contact Dell support tomorrow.
Could you recommend a way to contact Dell support for Isilon products?
We are located in Hungary.
Thank you!
Andrew
DELL-Josh Cr
Community Manager
•
9.7K Posts
•
43.3K Points
0
January 4th, 2021 10:00
To replace the boot drives you are going to need to call phone support as they will have to walk you through the process. It looks like they are not reporting the health correctly.
DELL-Josh Cr
Community Manager
•
9.7K Posts
•
43.3K Points
0
January 4th, 2021 11:00
Here is the contact page for Hungary https://dell.to/3rWgFH8
Phone Number: 068 0016172