Announcement Banner
UNSOLVED

pvenini

updated

11 years ago

P

pvenini

1 Rookie

•

3 Posts

1

3073

January 14th, 2016 10:00

Problem with VNX5200

Hi, my name is Pablo Venini and I’m CTO of Mercado de Valores de Rosario S.A., an stock exchange based in Rosario, Argentina. We bought  a VNX5200 array as part of a project to modernize our IT infraestructure, along with high performance servers, to support our expansion plans.  We choose EMC over its competitors because we needed an array capable of surviving failures  (we ran a trading system that is connected to the nationwide trading infrastructure, and monitored by government enforcement agencies); and the VNX 5200 offered dual active/active controllers, mirrored drives, redundant power sources and interfaces and remote proactive monitoring (we acquired a premium support offering from EMC). This array was installed in an APC rack in a  datacenter with dual online UPS, dual diesel generators, dual redundant precision AC and environment monitoring.

We had an initial array that after some months of being installed started to show an overtemperature warning that led to the array being shutdown and restarted immediately; this happened every five minutes all day and prevented our servers from working with the data stored in the array and eventually destroyed all its disks due to the tear and wear related to this situation. This failure was never explained nor resolved, even after the whole chassis and its parts were replaced (the overtemperature alert was verified by the onsite service representative in situations when our datacenter was less than 20°C ambient temperature). 

We were offered a mechanical replacement that was installed in the same site and worked well for a few months until on 12-14-2015 the array gave an overtemperature warning and shutdown itself and immediately restarted with half it LUN’s down (neither the ambient monitoring system gave an alert nor any other equipment in the datacenter had any thermal alert, we also have an emergency shutdown system that works on overtemperature conditions and it didn’t trip), this problem was solved by technical support and the array worked for a week until on 12-21-2015 the array went down completely, taking out our whole infraestructure in the middle of the trading session.

Technical support connected remotely and could only manage to bring it online after 5 hours of working with it, only for the array to fail again 8 hours later. This prompted another lengthy session with technical support and another failure 8 hours later. What technical support found was that 1 system disk and 2 data disks where showing high temperatures (65°C), this overtemperature prompted the array to shutdown itself and the prevented it from booting again. Strangely, the disks sitting next to the overheating disks were showing a temperature of 20°C.  This was again verified by the local representative when the ambient temperature was below 20°C. Technical support directed for the disks to be replaced, and only after replacing the system disk that had failures the array stopped shutting down itself (this happened on 12-25-2015).

All this situation is complicated by a lot of administrative problems we have had with EMC, which includes:

-The first array was wrongly associated with a different firm (Mercado a Termino de Rosario)

-All our EMC equipment and services where assigned to different administrative sites, even though they are on the same place

-The ESRS we had assigned in our administrative site in fact was from another firm, so these arrays were never  proactively monitored (this was discovered a week ago)

-The second array wasn’t registered properly, so we had to open all our service requests under the ID of the first array (which was never deleted from our site)

Due to the instability of the array we were forced to move all our workloads to an alternative environment (an environment that lacks the processing power we had in the previous one and is not suitable for extended operations), and we have been in this situation since 12-21-2015 as we are hesitant to use the array again since we have no more confidence on it. Right now we are waiting for a concrete explanation on behalf of EMC as to why two arrays failed the same way, and we also want to know why an equipment that has redundancy in all its parts fails catastrophically when one of its components fails (a system disk that is mirrored and used by a service processor that itself is also redundant).

In case we don't receive a quick and satisfactory answer we are going to return this equipment in exchange of the money we paid, and claim damages for all the inconveniences our firm suffered.

Pablo Venini

CTO

Mercado de Valores de Rosario

pvenini@mervaros.com.ar

54 0341 4210125