Unsolved
This post is more than 5 years old
3 Posts
0
3528
February 27th, 2014 05:00
SP failover takes 2 minutes
Hi,
I just implemented a VNXe3300 running iSCSI server for VMFS datastores for VMware + Fileserver running for CIFS shares.
We are testing availability and notice during a controller reboot we do from mainteance in Unisphere it takes up to 2 minutes before the iSCSI server and corresponding datastores/targets for VMware are available again. Of course VMware doesn't like this long of a timeout and I can't imagine this is HA as the VNXe series is meant to offer.
We got both controllers hooked up (crossed) with 10B Fiberchannel cables & 10GB Switches and from switch level everything works as a charme, we tested switch outage which works fine but from Storage Processor level running iSCSI server it behaves not High Available.
Also the management port itself takes this long to get available again.
Is there anything we can do to speed up the behavior of SP failover to prevent VM's to go down.
Many thanks for your thoughts..
Regards,
Marcel.


chrislogo
98 Posts
0
February 27th, 2014 08:00
Hello Marcel,
I have few questions on your configuration
Is the ISCSI and CIFS Servers on different SP's ?
What is the IP range you are using for the iSCSI and the CIFS Servers and are they routable with each other ?
Chris
mvmilt
3 Posts
0
February 27th, 2014 13:00
Hi Chris,
CIFS we have not configured yet, but is there a way to get CIFs and iSCSI active on separated SP's? I thought only 1 SP is active at 1 time so all servers you create are only at 1 SP visible..?
If there is a way to create it on 2 SP's i'm suspecting that is why it takes so long, as that is not done right nowI was looking for a way to get the iSCSI server added to both SP's.
When I configure the iSCSI server I can only select the ports on the PS that are on that SP, I can't see the SPB's eth ports. I was looking for a document on this but so far have not found it. If you got any steps describing this please let me know.
I was thinking you would need to:
1. Configure on SPA the iSCSI server
2. Add the IP on 1st eth FC ports we use on SPA
3. Add a 2nd IP on 2nd eth FC port we use on SPA
4. Create the VMFS folders (we use 3)
5. Select the ESX servers to access the VMFS datastores
(so far we did already)
This works but gives when SP is booted a 1,5 to 2 min unavailability of the VMFS stores which I can't think is good..
Shouldn't we repeat above after reboot of SPA to SPB and select there the eth ports on the then active SPB and while creating a 2nd iSCSI server?
Thanks in advance,
Regards
Marcel.
chrislogo
98 Posts
0
February 28th, 2014 10:00
Hello Marcel,
VNXe SPs work active/active in the sense that they both actively serve data so yes you can configure CIFS on SPA and iSCSI on SPB.
An important requirement for the failover to work is that the the physical connections for both the SPs should be identical/symmetrical.
What I mean is eth2 on SPA should be connected to the same network as eth2 on SPB for the failover to work correclty.
When you configure an iSCSI Server on SPA you can only use eth ports from SPA , the ports on SPB will automatically be used on a failover if they are symmetrical
Also VNXe does not use FC ports but normal RJ45 LAN ports for iSCSI over LAN.
Attaching the HA document which can help understand the failover configurations
Chris
chrislogo
98 Posts
0
February 28th, 2014 10:00
VNXe HA Document attached for your reference
1 Attachment
docu35554_White-Paper_-EMC-VNXe-High-Availability-Overview.pdf
Leo Li
4 Apprentice
•
9K Posts
0
March 6th, 2014 01:00
Hey,
VNXe3300 can support "Two-port 10-Gb/s optical iSCSI I/O module". I'm not sure whether it can speed up the failover.
Just checked that Two-port 10-Gb/s optical iSCSI I/O module will not help speed up the failover it can only help increase the access performance. Hope others can provide a good solution for you!
dcom8
2 Posts
1
March 6th, 2014 14:00
It is my belief that the VNXe will take that long to fail the Primary SP to the peer SP during an SP outage regardless. We are seeing the same 2min times.
If it is just a path outage and there are multiple paths configured and available then the host software should immediately find the next available path. Which you state works fine.
If the management port goes down the two min delay will take place to fail the management NIC to the peer side as well is normal.
You must remember that the VNXe series is the "entry/essential" line of EMC NAS products and has its inherent limitations.
Part of the issue is that the "data movers" and "SP" are one in the same in the VNXe series where they are not in the parent VNX series where data mover fail-over is relatively fast on the VNX and older Celerra models.
mvmilt
3 Posts
0
March 6th, 2014 23:00
Hi,
Here an update of where we are currently with the VNXe3300 for this customer.
The VNXe3300 is a good product in terms of high availability. 2 active SP's are usefull to balance things if you like but do not make the running part on 1 SP high available without downtime in case of SP failure.
The FSN mechanism works great, you pull out eth11 of SPA and eth11 of SPB continues, you pull the second eth21 of SPA also out and eth21 of SPB continues with 1 timeout in ping. This is working excellent.
What I personally find disappointing on this VNXe3300 is the management port. This only runs active on 1 port and does not continue to the other port of the other SP when you unplug the cable. You loose the management and it takes about 10 minutes before it allows you to get back in the console/webgui (and then is active from the other SP it's mgmt. port).
We further tested failover after reboot of an SPA to B or vice versa:
The 3 VMFS VMware Datastores are not accessible for at least 30 seconds minimum.
If you look at them in vsphere you see them italic and coming back after about 30-40 seconds.
The fact it's "gone for 30 seconds" seems something we can't change, the fact the Data Movers and the SP are in the same on the VNXe seems to cause it to take longer to become accessible when its moved to another SP in case of failure of the SP or with reboots (firmware upgrades)
You can tune your hart out at iSCSI protocol level but if the network connection is not able to ping and it's gone for 1/2 a minute you won't change anything to that.. you can only make sure the app-side doesn't respond to it with declaring the path dead etc.
If somebody has VMware knowledge in relation to the VNXe I really like to know what the best timeout values would be from the iSCSI adapter in Vsphere on the ESX nodes, I though about setting LoginTimeout to 60 seconds and maybe disable the delayed ack? Any suggestions?
The thing we are worried about is the sustainability of the VM's after moving them with Storage VMotion and getting an outage. 1 VM now stays up and stalls writes when it happens, but does not goes down. But what happens when we have 20 server VM's with Exchange included running. Will it then still keep up when being gone for 30 seconds..
Many thanks
Marcel
Shivanand1
70 Posts
0
March 12th, 2014 23:00
Hi Marcel,
Failover of services from one to other SP would take around 3-4 seconds. Usually for cifs & nfs users this short interuption will go unnoticed. However for iSCSI hosts (VMware datastore, generic iSCSI lun) even this short interuption might bring down the system (VM).
Hence it is recommended to increase the timeout value on all the iSCSI hosts to set as "60" seconds.
Below is the procedure to increase time you value for ESX hosts,
1.Log in to the vSphere Client and select a host
2.Navigate to the "Configuration" tab
3.Select "Storage Adapters"
4.Select the iSCSI vmhba to be modified (typically the iSCSI software initiator)
5.Click "Properties"
6.Select the "Dynamic Discovery" tab
7.Select the IP address for the Equallogic group
8.Click "Settings"
9.Click "Advanced"
10.Scroll to "LoginTimeout" and set the value to 60
11.Repeat 1-10 for all applicable hosts/servers
12.Host reboot is necessary to apply changes
Regards
Shivanand