WRT VNXe 3100 configuration, and link failover and such, here’s what I don’t quite understand:
The docs refer to the two storage processors on the system as SP’s. This is terminology held over from CLARiiON, and strongly implies to me that the SP’s run active/active. But I also know that there is something analogous to DART under the covers, allocating iSCSI LUNs and NAS shares/exports.
So, how does this work? Is this DART-like code running on just one SP at a time? Both? If both, then what? What does that mean from a networking and storage access perspective? All LUNs and NAS shares/exports available on all configured network ports?
I don't understand, and I haven't found a document that explains it.
While the VNXe does not run DART, it does act similarly to DART and shares a lot of functionality with Celerra/DART.
For network HA the VNXe acts a bit differently from Celerra though. Since Celerra used an Active/Passive model, you would build network redundancy within each datamover to handle network failures, and then duplicate the config on the passive datamover in the event of datamover failure. In Celerra configurations, this usually ended up employing the FailSafe Network and resulted in some ports being designated as passive which reduced overall bandwidth available to the datamover. VNXe is different due to its Active/Active model.
“For network high-availability features to work, the cable on each SP needs to have the same connectivity. If Port 0 on SPA is plugged in to Subnet X, Port 0 on SPB must also be plugged in to Subnet X. This is necessary for both server and network failover. If a VNXe server is configured to use a port that is not connected on the peer SP, an alert is generated. Unisphere does not verify if they are plugged in to the same subnet, but they should be, for proper failover. If you configure a server on a port that has no cable or connectivity, the traffic is routed over an SP interconnect path to the same port on the peer SP (just a single network connection to the entire system is not recommended).”
--Page 30, EMC VNXe Series Storage Systems, A Detailed Review
“Network paths – The VNXe supports network pass-through to provide network path redundancy. If a network path becomes unavailable due to a failed NIC, switch port, or bad cable, network traffic is re-routed through the peer SP using an inter-SP network, and all the network connections remain active.”
--Page 31, EMC VNXe Series Storage Systems, A Detailed Review
With this in mind, I personally would create an LACP group of all ports on SPA connected to one switch and an LACP group of all ports on SPB connected to a different switch. That way you get the benefit of ALL ports for traffic, network switch redundancy, and SP redundancy.
Description of SP Active/Active failover…
“When an SP experiences hardware or software failure, reboots, or a user places it in Service Mode, a failover occurs. Storage servers that use the out-of-service SP fail over to the other SP if it is available, with minimal disruption between the VNXe system and connected hosts.
…
After the SP is available again, the storage servers fail back to the original SP. The failback policy can be configured through Unisphere.”
So, if I'm to understand this correctly, you're saying that the SP's run active/active like that of the CLARiiON, but the iSCSI and NAS provisioning functionality is controlled by something analogous to DART?
This surprises me, because only 6 of the 12 network interfaces are showing up as available on the particular VNXe 3100 implementation that we're currently doing. It's very Celerra-esque in how it is configured.
Is there a document that really does a good job of describing all of this. It is not clear to us at all, and we've been to all the VNXe implementation training
VNXe is active/active in the sense that both SPs service requests at the same time. Any particular LUN or Share will be owned by one or the other SPs under normal circumstances. If one SP fails, then the other SP takes over all of the volumes/shares from the failed SP. This is all very similar to the way Clariion SPs handle LUN failover.
The interesting/cool thing about the VNXe’s hardware is that there is an internal network switch between the SPs, so if a client loses access to the owner SP through the network, but the SP is still online, it can send requests to the secondary SP which will forward the requests internally through the internal switch.
Richard, thank you for your complete, and authoritative answer.
Just a suggestion... There's much to be interpreted from all of the information presented in that paper. Somehow, a higher-level description, sort of like a section entitled "VNXe Networking for Dummies" if you will, needs to be put into that paper. Your narrative is a very good start on that.
I have a 3300 that I am implementing now but would like to utilize the 10gig ports and have some form of I/O load balancing.
What I was thinking was utilizing both 10gig interfaces on SPA on a specific network and connect them to a dedicated switch. Then configure both 10gig interfaces on the SPB on a separate network and connect them to a separate dedicated switch? My vmware hsots would have dedicated vmkernels/pnics on both networks for failover and redundancy. I would like to utilize both paths for some form of load balancing but I am not quite clear on how to achieve this with the vnxe.
I worked with EMC support on configuring multipathing with my vnxe3300. I am using (2) cisco 4948 and (4) 10gig ports on the vnxe.
Its configured like so:
SPA
eth10 192.168.55.5 -
> 4948A
eth11 192.168.66.5-------> 4948B
SPB
eth10 192.168.55.6 -
> 4948A
eth11 192.168.66.6 -
> 4948B
I then have all datastores in vmware configured for round robin multipathing. However I am seeing a fair amount of different paths (missing/broken/dead) between all my hosts and the vnxe3300. Waiting to hear back from EMC as to why.
However even without the transparency its been working quite well for me at this point. Its almost two basic though, I want to see information but its just not given to me from EMC and unisphere.
I'm reading through this thread and appreciate all of the detailed info.
I agree with Eric though, EMC needs to be more transparent about how this is supposed to be configured. What is the best practice for configuring multipathing on a VNXe? Some people say it is to split the SP ports into different subnets. Port 0 of each SP is connected to physical switch 0 and is in subnet X. Port 1 of each SP is connected to physical switch 1 and is in subnet Y. You mentioned configuring a LACP group with all ports on each SP and connecting them to the same physical switch. Which is the correct way (I'm asking specifically for a VMware use case).
In the VNX techbook for VMware on page 50 it describes connecting the ports on each SP to different switches like I described above, but using the same subnet for all ports. I know CLARiiON arrays pre-FLARE30 had a bug that required using different subnets, which is how I got used to doing things. Has that changed with VNX and VNXe?
Total assumption on my part, but I would tend to believe its AULA (whether its referred to as AULA I dont know) because EMC had me configure the ds for round robin and being that EMC owns vmware I would hope they would be the ones to know.
It has been stated that the VNXe is active/active. Even though each LUN is "owned" by a specific SP, there is inter-SP connectivity inside the array that allows IO to be passed from the "non-owning" SP to the owning SP. This sounds an awful lot like ALUA, which is found in the CX and the VNX devices, but I've never heard of anyone calling this ALUA on a VNXe. Is the VNXe ALUA or not?
If it is ALUA, then round robin would be an appropriate PSP for VMware. If it ISN'T ALUA, then I would think that MRU would be the appripriate policy.
The VMware HCL lists the VNXe3100 and 3300 as being supported with MRU as the PSP. However there is a disclaimer right above it that says some vendors might support RR, but you need to contact them for info.
It may help everyone here to understand that VNXe networking is NOT the same as VNX networking. VNX follows the same design guidelines as Clariion/Celerra.
With FC/iSCSI block on Clariion and VNX, ALUA is available to redirect IO to the alternate (non-owning) SP. Those redirected IOs are processed by that secondary SP and then passed to the owning SP, so the IO is actually being handled by both SPs in certain ways. This is why ALUA typically directs IOs to the owning SP if it can.
With NAS on Celerra and VNX, the front end datamovers are active/passive, each active datamover owns some set of filesystems/shares/iSCSI LUNs, etc and there is a passive (not fully booted) datamover that takes on the complete identity of a failed datamover when needed.
VNXe is not the same as Clariion, Celerra, or VNX, it is actually a newer design. It is Active/Active in that both controllers are actively processing IO simultaneously. However, there is still a notion of ownership like Clariion in that the filesystems/LUNs are owned by one of the two storage processors. If IO is directed to a port on the non-owning controller, there is actually an internal Ethernet switch of sorts that forwards the traffic. It is not ALUA. That said, I would not recommend directing IO to the non-owning controller under normal operating scenarios as it may affect performance. While you could follow the same network guidelines as VNX NAS for a VNXe, the fact that VNXe forwards IO between controllers means you have more options than a VNX as far as leveraging LACP and still getting switch level redundancy. It also depends on whether you are using NAS or iSCSI. With NAS you could easily connect all of the ports on one controller to the same switch and put them in a single LACP group. Then connect the second controller to a second switch/second LACP group, and use the internal switch/forwarding to handle switch failures. With iSCSI you may want to have two “fabrics” in which case you would want to mesh the controller connectivity into both switches using two ports (or two LACP groups) per controller.
Storagesavvy
474 Posts
3288
0
Posted June 23rd, 2011 13:00
While the VNXe does not run DART, it does act similarly to DART and shares a lot of functionality with Celerra/DART.
For network HA the VNXe acts a bit differently from Celerra though. Since Celerra used an Active/Passive model, you would build network redundancy within each datamover to handle network failures, and then duplicate the config on the passive datamover in the event of datamover failure. In Celerra configurations, this usually ended up employing the FailSafe Network and resulted in some ports being designated as passive which reduced overall bandwidth available to the datamover. VNXe is different due to its Active/Active model.
“For network high-availability features to work, the cable on each SP needs to have the same connectivity. If Port 0 on SPA is plugged in to Subnet X, Port 0 on SPB must also be plugged in to Subnet X. This is necessary for both server and network failover. If a VNXe server is configured to use a port that is not connected on the peer SP, an alert is generated. Unisphere does not verify if they are plugged in to the same subnet, but they should be, for proper failover. If you configure a server on a port that has no cable or connectivity, the traffic is routed over an SP interconnect path to the same port on the peer SP (just a single network connection to the entire system is not recommended).”
--Page 30, EMC VNXe Series Storage Systems, A Detailed Review
“Network paths – The VNXe supports network pass-through to provide network path redundancy. If a network path becomes unavailable due to a failed NIC, switch port, or bad cable, network traffic is re-routed through the peer SP using an inter-SP network, and all the network connections remain active.”
--Page 31, EMC VNXe Series Storage Systems, A Detailed Review
With this in mind, I personally would create an LACP group of all ports on SPA connected to one switch and an LACP group of all ports on SPB connected to a different switch. That way you get the benefit of ALL ports for traffic, network switch redundancy, and SP redundancy.
Description of SP Active/Active failover…
“When an SP experiences hardware or software failure, reboots, or a user places it in Service Mode, a failover occurs. Storage servers that use the out-of-service SP fail over to the other SP if it is available, with minimal disruption between the VNXe system and connected hosts.
…
After the SP is available again, the storage servers fail back to the original SP. The failback policy can be configured through Unisphere.”
Richard J Anderson
1 Attachment
h8178-vnxe-storage-systems-wp.pdf
h8178-vnxe-storage-systems-wp.pdf