The last two weeks we had about 3 times 30 seconds of downtime caused by out switches. The problem is that a lot of packets are being discared. These packets are received on our core switch. We have configured two 6224's in a stack connected to multiple 3500's. They are connected via two cables and configured as port channel on both sides.
The regular port channel does not show any errors:
Unit 2 discarded some frames! I see this behavior on all switches connected to the second unit.
Port configuration:
CoreSwitch#show running-config interface ethernet 1/g8 channel-group 8 mode auto
CoreSwitch#show running-config interface ethernet 2/g8 channel-group 8 mode auto
CoreSwitch#show running-config interface port-channel 8 service-policy in VOIP spanning-tree guard root switchport mode general switchport general allowed vlan add 8,10,12,20,120,200 tagged switchport general allowed vlan add 1 tagged
AccessSwitch#
interface port-channel 1 switchport mode general switchport general allowed vlan add 8,10,12,20,120,200 tagged exit
interface ethernet g3 switchport mode general channel-group 1 mode auto
interface ethernet g4 switchport mode general channel-group 1 mode auto
A soon as we lost connection I rapidly connected to the switch and did a show process cpu command to see what what going on. The cpu utilisation is higher than normale but not unacceptable high. Normally it's below 10% and now the 1 min timer raised to 32%. 18% was caused by the ipMapForwardingTask, which might indicate a buffer issue?
Any troubleshoot tips? Vlans are the same. The stranged thing is that only unit2 is discarding packets. Vlans are configured the same so that doens't explain the discards. We also notice a connection loss for a short time so that's not the issue.
ipMapForwardingTask is related to the IP stack, so typically when it is high there is too much traffic for the switch to keep up. Is the switch firmware up to date? Is spanning tree enabled? If it is switch 2 the spanning-tree root switch?
The switch is running at almost the latest firmware version (3.3.11.2) so I would not expect any fixes for this issue regarding the release notes on the latest version (3.3.12.1). De stack is the root switch, configured with the spanning-tree priority 4096 command. Spanning-tree looks normal for me. For example, channel 3 is connected to 1/g3 and 2/g3 and so on. The only exception for this rule is po1, which is bound to 1/g21 and 2/g22. These ports are not connected via port channel:
Good point. The logs are still in there so I've taken a look. There are some intrestring lines:
<189> MAR 27 10:11:14 192.168.20.1-2 TRAPMGR[124132592]: traputil.c(611) 4455 %% Spanning Tree Topology Change: 0, Unit: 1
<189> MAR 27 10:11:18 192.168.20.1-2 TRAPMGR[124132592]: traputil.c(611) 4456 %% Spanning Tree Topology Change: 0, Unit: 1
<189> MAR 27 10:11:22 192.168.20.1-2 TRAPMGR[124132592]: traputil.c(611) 4457 %% Spanning Tree Topology Change: 0, Unit: 1
<189> MAR 27 10:11:26 192.168.20.1-2 TRAPMGR[124132592]: traputil.c(611) 4458 %% Spanning Tree Topology Change: 0, Unit: 1
<189> MAR 27 10:11:30 192.168.20.1-2 TRAPMGR[124132592]: traputil.c(611) 4459 %% Spanning Tree Topology Change: 0, Unit: 1
<189> MAR 27 10:11:34 192.168.20.1-2 TRAPMGR[124132592]: traputil.c(611) 4460 %% Spanning Tree Topology Change: 0, Unit: 1
<189> MAR 27 10:11:37 192.168.20.1-2 TRAPMGR[124132592]: traputil.c(611) 4461 %% Spanning Tree Topology Change: 0, Unit: 1
<189> MAR 27 10:11:41 192.168.20.1-2 TRAPMGR[124132592]: traputil.c(611) 4462 %% Spanning Tree Topology Change: 0, Unit: 1
<189> MAR 27 10:11:45 192.168.20.1-2 TRAPMGR[124132592]: traputil.c(611) 4463 %% Spanning Tree Topology Change: 0, Unit: 1
<189> MAR 27 10:11:49 192.168.20.1-2 TRAPMGR[124132592]: traputil.c(611) 4464 %% Spanning Tree Topology Change: 0, Unit: 1
I see this kind of messages pretty frequently. So the problem might be something in the spannning tree topology again... I am going to do some research on this one and let you know the results.
I am not sure anymore if spanning tree is the problem. A topogy change will occur even when a client turns on so these messages are very common (or is it only showing local topology changes?). Also, most ports are configured with the spanning tree root command, so STP root switch changes should be very rare or even never happen....
All the switches are configured with RSTP.
Because the switch is configured as stack, there is no need to configure STP on both switches? (it's even not possible I guess).
The only strange thing I see so far is that one access switch is connected via three cables (two UTP one fiber). The fiber is not in use anymore (administrative shutdown) but to make sure it doesn't harm anything I'll remove the cable tommorow. Despite of the admin shutdown the port is still up on other side.
DELL-Josh Cr
Community Manager
•
9745 Posts
•
44407 Points
4541
0
Posted March 27th, 2015 12:00
Hi,
ipMapForwardingTask is related to the IP stack, so typically when it is high there is too much traffic for the switch to keep up. Is the switch firmware up to date? Is spanning tree enabled? If it is switch 2 the spanning-tree root switch?