UNSOLVED

IT-girl

updated

12 years ago

IG

IT-girl

4 Posts

0

10117

September 24th, 2014 13:00

throttling replication offsite.

We have 2 EQ, PS6500 that  have been in place for 3 years.  We have 7 volumes, 3 are set up into a replication group for replication to our offsite location, replication kicks off hourly.   We have a 100mb metro-e connection to transfer the data.  The EQ is on it's own vlan and we have an existing throttling policy that should be throttling the traffic on that vlan and port 3260 to 80mb.  That way we do not saturate the pipe.  Generally our replications are approx. 10gb per hour per datastore (for a total of about 30GB).  Twice a day we have larger transfers on one of the datastores, these deltas can be 100-200gb.   In those instances, Replication takes longer than 60 mins and will take a couple hours to get caught back up.

What we have discovered is when all 3 replications kick off at the top of the hour, the throttle fails and the metro-e is saturated at 97%.  This also causes the router CPU to spike, causing latency and jitter which affects the SIP/VOIP traffic that is on a different VLan and different network pipe (30mb mpls).

If all 3 remain running, the pipe continues to run at 97% until the first stream completes, then the throttle kicks back in and the remaining 2 datastores continue to run at or below the 80mb throttle.  (sometimes we see it staying solid at 65, other times 80).

Once we spike to 97%, I can pause the replication of 1 or 2 datastores, this will bring the traffic back down to 80.  I usually can then add back in 1 or both of the other datastores and it will remain in the policy.  Although, for unknown reasons, suddenly it will spike back up to 97% midway thru the replication and the only way to back off is to pause replication on one or two datastores then resume them incrementally.

This is having a critical impact on our production environment as well as our phone/call center because of the voice issues that are affected when the traffic spikes. 

I'm hoping there might be others that are running in similar environments and would like suggestions on how you handle replicating your data across.  How much data, what size connection and how long does it take?