Unity 650F LUN latency sometimes ramps up to hundreds of milliseconds starting at 5:03 PM and ending around 6PM. Happens at random, but seems to occur more often when we're moving a lot of data around, creating and deleting large LUNs. Drive bandwidth becomes VERY high but CPU and front-end activity remains normal.
Latency example:
Drive bandwidth:
LUN IOPS does not follow a similar pattern:
LUN bandwidth actually drops because the system is performing so poorly:
CPU remains normal:
EMC\C4Core\log\c4_safe_ktrace.log contains lines similar to this during the period of high latency:
Drive bandwidth from 6PM to 7PM. Note the peculiar pattern. "Packs" of disks seem to be synchronizing or rebalancing. We do not see this access pattern under any other occasion. It only shows up during these 5PM "storms."
We've opened tickets twice. The first time, the result was "it has to be something on your side generating this activity" when it's very clear that front-end activity was business-as-usual during the 'IOPS storm.'
The second time resulted in us receiving an "everything mostly looks fine" report from support that was a mishmash of incohesive graphs and observations and it was difficult to tell exactly what we were looking at. The case was submitted with attachments from internal logs (c4_safe_ktrace and C4Core\log) showing clearly a RAID group rebalance getting underway at the exact time of the IOPS storm. They ignored it.
DELL-Sam L
Community Manager
•
8056 Posts
•
34380 Points
1105
0
Posted April 16th, 2021 16:00
Hello richardm112,
If you still have the service request# can, you please send me them via private message & I can try to get them escalated & reviewed.