
UNSOLVED
D
drominore1
1 Message
0
562
October 7th, 2014 06:00
VNXe3100 upgrade woes
We recently had a CE do a remote upgrade from V2.4.1 to 2.4.3. During this upgrade which we had expected to be zero downtime the unit disconnected for about 20-30 minutes and caused significant impact. The upgrade did finish and was successful but I requested a Reason For Outtage review and sent service logs taken just after all was back up. One engineer called back and relayed that since we were a Hyper-V shop (2008R2 converting to 2012R2) and that the hypers were all using CSV on the VNXe, that MS has a timeout issue where it will appear to lose communications and stall all the guests on the SAN. He also mentioned that if we were a VMWare shop he could adjust those timeouts to reduce the chance of the issue to about 5%, but in Hyper-V its about a 50-50 shot that you will lose communications during an SP switchover.
I have areal problem with this explanation and I need some kind of documented backup of it from either EMC or MS. I have requested such in my RFO but at this time that has not been completed. Has anyone else experienced this issue? Also of interest is that I was able to remote back in and get to the web interface, which instead of having the portal site had a placeholder webpage. This strongly suggests to me that this was NOT an MS issue but rather one where BOTH SP's rebooted (or at least the one with the upgrade in progress was visible to outside). The error logs visible through the portal though do NOT show a crash, just nominal upgrade messages.
I need to answer to my CTO on what happened here and its likeliness to occur again. Our luck thus far with upgrades are 2 successful, two failed at this time which matches the ratio the engineer mentioned. What will happen if the unit itself decides through internal checks to switch SP's? Are we expected to carry a 50-50 chance that this automated function would take down our entire storage array, probably at the most inopportune times?
Responses (0)
Solutions (0)
