Replication warnings in Celerra Manager after DR move
We had 2 NS20s on site. One was production, the other for DR/replication. When the replications were configured, both NS20s were on the same LAN and everyting was fine.
We eventually moved the DR NS20 offsite to a data center. The sites are connected with a 20mb fiber WAN connection.
After the move, we started receiving the following warnings in Celerra Manager. The replications are still working and marked OK, but 20-30 warnings pop up daily about drops in the connection.
Severity: Warning
Brief Description: Slot 2: Primary=fs31_T1_LUN0_APM00081800540_0000_fs27_T1_LUN0_APM00080701610_0000(alias=iSCSI_LUN0_Replication), transferring. Data connection down. Retry.
Full Description: communication between source and destination is down.
Recommended Action: Bring network communication between source and destination up for the data to be transferred from source to destination.
The warnings are either single entries or up to exactly one minute to the second. They happen with all of the replications that we have setup.
After running bandwidth montioring, we are only using 6-7mb of our 20mb pipe.
Although I do think this is a network connectivity issue, our provider of the fiber line says everything is fine. Other tests do not show any drops in the connection. The Celerra Manager is the only device reporting connection drops.
Has anyone seen these warnings after moving replications from a LAN to WAN connection? Is there settings in the Celerra to adjust because of the change in bandwidth?
I understand. The clock skew they fixed does not have any relation with the error message you are seeing. It eventually block your sessions if it's greater than 10 minutes.
I suggest you to ask for a NAS code upgrade on both sides. As I mentioned, this might not "fix" your error messages, but will give you several fixes and enhancements.
I did open 2 tickets with EMC . The first ticket they found and fixed the following: Time on the destination control station is one hour behind. The interconnect time was skewed and this was fixed.
However, this did not fix the issue with connection drops. We have talked to our EMC rep and he said that we could run the Fiber Analyzer from EMC but we would be charged for the service. They did not recommend running this because it is probably an network issue with the connection.
I am in the middle at this point because EMC is telling me everything is fine and our consultant is saying everything is fine with line.
If you think it is a network problem and have noticed a time of day when it consistently happens, could you script a periodic/frequent ping to the IP address of the replication interface from a workstation to see if you lose connectivity? If the network goes down the ping should fail while it is down and you should be able to see it in the log if you capture the ping output. That'll give you one more data point.
Also, you might see if the server logs give you any indication of what is going on.
The warnings appear randomly and do not fall into certain periods of time. Sometimes the warnings appear in the early AM during non-productions hours and other times they happen during the day. The warning are mostly single entries but sometimes they appear for exactly one minute to the second.
At first, our consultant blamed the NS20 for saturating the WAN line and causing the warnings. They wanted us to throttle the NS20 to slow down the speed. After running some bandwidth monitoring, we found that wasn't the case.
The consultant is now opening tickets with the WAN provider and they have been running some tests. On Wednesday, they did some "repairs" on the line but the warnings are still appearing.
I'm trying to cover the bases with the NS20. I still do not know if this is 100% a WAN issue. I have a feeling the WAN provider will come back and say the line is fine, which it is to an extent. The replications are working and it is the small drops causing the Celerra to throw a warning.
Is it possible to configure the Celerra to be less restrictive on the warnings it displays or should I still continue to try to find the source of the issue?
I have this issue for awhile now. replicating CIFs and iSCSI. Out of 12 replications I have an issue with 3 of them.
Latest code 5.6.45-5 does NOT resolve the problem. Support suggested to looking to my network after reviewing packet capture
which indicates a lot of "TCP Previous Segment Lost", "TCP Dup Ack", TCP Out-of-Order" captures.
Can anyone take a capture before your firewall and check if you see the same?
Weird thing is it is not happening with all replications so if it would be networking issue I would assume all my sessions would have the same problem.
gbarretoxx1
2 Intern
•
366 Posts
754
0
Posted October 2nd, 2009 09:00
What's the NAS code level running on source and destination boxes ?
What's the latency on this link ?
If you do a server_ping server_2 what's the response time ?
Gustavo Barreto.