UNSOLVED

brettesinclair

updated

11 years ago

B

brettesinclair

2 Intern

715 Posts

0

3002

January 16th, 2016 22:00

High incidence of node failure

Hello everyone,

We have 2 Centera clusters at remote sites that over the last 6 months have seen high occurrences of node failure.

These can be access or storage nodes and generally they can just go offline at random, requiring disk transplants into new node chassis. At the same time they will often trip the left power rails for the rack. (not a big deal with dual feeds).

For example on one 16 node cluster we have had  7 node failures in 6 months, and similar stats on another cluster.

I have raised it with our TAM who advised that they had identified a hardware fault on a particular node version and were just replacing them

"as they die", rather than a co-ordinated replacement. This is in line with what the onsite tech told me, who advised this has kept the local FSS's busy in that time.

The replacement nodes have a Model # 100-580-573 A15 and the nodes that are failing are 100-580-711-01 A02.

Anyone seeing similar things in the field ?