
UNSOLVED
PS6510E latency spikes
Equipment:
1 PS group: 2 x PS6510E, 1 x PS6210S, 4 x PS6210E all at v7.1.4
PowerEdge M610,M910 blades w/ 10G intel NIC
m8024k blade switches, 8024F san switches
Software:
ESXi 5.1 U3
Windows 2008R2
Red Hat Enterprise 6.6
For a while now, been seeing some latency spikes (mainly write) on our PS6510E arrays. Basically what happens is everything is going well, then all of a sudden we get a latency spike on just one 6510E member. (The 6510Es are each in their own pool) Most of the time its just write latency, and read latency remains low. This is observed from SANHQ, but our VMs that have datastores on these arrays exhibit a high 1minute load average that coincides with the spike. ESXi hosts log messages similar to "lost access to volume... due to connectivity issues" followed a couple seconds later by "Successfully restored access to volume.. following connectivity issues". This is logged for nearly all the datastores on the particular member (and only that member). It recovers almost immediately, but the interruption is enough to cause a brief load spike on the VMs. The initiators do not seem to be logging out and back in during this time.
Steps taken: I have done my best to configure everything according to best practice. (Disable delayed ack, LRO, using MEM 1.2) and verified this is all set correctly. Jumbo frames on, not using default vlan, etc.
I have not noticed this on our SSD 6210, or any of the other 6210s but those are also not being used heavily for vmware right now. Also I have offloaded a lot of our "heavier" IOPS volumes to our newer 6210 arrays, to reduce load on the 6510s, but this has not seemed to help all that much. You can see below the array doesn't seem to be that busy. I have the line hovered over the spike:
Thanks for ideas/suggestions.
Responses (0)
Solutions (0)
