Unsolved
This post is more than 5 years old
3 Posts
0
24530
January 6th, 2012 01:00
M6220 Stack member not responding to management ip after failover.
Hi,
I have a problem with my new M6220 switches that are placed in my new M1000e enclosure (all with the latest firmware.)
Two of my M6220 switches (placed in A1 and A2 of the enclosure) are in a stack connected by two stack cables from the stack modules.
When I initiate a failover by software or by removing the master from the enclosure, the stack member becomes the new master of the stack. The only problem is that the new master does not respond on its management ip address. The Chassis management controller shows me that the new master received the ip address of the old master but it just won't respond...
After that, to move back to the first master, I have to power cycle the second switch. But when I do that the whole enclosure and the blade servers are not responding for about a minute. after the some time everything is working as it should again.
Anyone an idea? Right now this is not a sytem that I want to take in production...
Thanks, Martijn.


seventhsven
19 Posts
0
January 13th, 2012 07:00
When you say "no response", do you mean that you can actually connect to the management (TCP 3-way handshake), but it does not return any data? Or does it time out completely, possibly not even answering to ARP requests for the management IP? Check with Wireshark or tcpdump, if you can.
We have hundreds of M6220 switches, and we have had numerous problems (several tens over the last few years) with the management becoming unavailable. It seems to happen unprovokedly after some random time (many weeks, typically), without any failover or similar to trigger it. Sometimes it is triggered by something like "show running-config" after the switch has been up for a few weeks or so as well.
The answers I've gotten from Dell support is to upgrade to the latest firmware to see how that fares. We've run 4.1.1.9 for a shorter while, so I don't think we've hit the issue yet, but I'll check it. There are no hints about how this problem can be investigated once it hits. When even a physical serial console gives zero activity, I don't know what more to try.
I'm not sure your problem is the same as ours, but I'd thought I'd at least comment with the info I have.
As for redundancy when stacking the two blade switches together: My experience is that with the flaky operation of these switches, it is a bad idea. An alternative, if you have multiple m1000e enclosures, is to stack the switches together in such a way that enclosure 1 switch A is stacked with enclosure 2 switch A, and enclosure 1 switch B is stacked with enclasure 2 switch B. Then configure the blades to use bonding with active/backup failovermode. That will give you stack redundancy too, as long as you can live with active/backup bonding.
seventhsven
19 Posts
0
January 17th, 2012 00:00
Okay, I see. Please notify this thread if you run into any problems.
Where did you get firmware 4.2.0.4a28, by the way? It's not available in the firmware directory on ftp.dell.com, and the only hit on Google is to this thread. Could you perhaps upload it somewhere (together with the release notes)?
FFriet
3 Posts
0
January 17th, 2012 00:00
We had to take the system in production so I'm not able to test it anymore... I'll keep my fingers crossed.
B.t.w. our firmware version is 4.2.0.4a28 and we only have one m1000e enclosure.
seventhsven
19 Posts
0
January 17th, 2012 00:00
Sorry, I found 4.2.0.4a28 through support.dell.com, direct download at downloads.dell.com/.../PCM6220v4.2.0.4_a28.zip .
What or earth, Dell? I wish they would have some consistency in their downloads. A well-thought out FTP or HTTP file serve with clever directory layout would be so much more useful than to navigate the support pages.