I am currently working on reducing high CPU utilization on one of our CX4-960 with 600 spindles. Both the SPs used to process 24000 front ends IOPS & the CPU used to be running at 80-90% each SP. We then bought in VNX & migrated complete Citrix farm which help reduce the work load & the IOPS drilled down to 12000 on CX4-960 with CPU running at 55-60 each SP. But our client wants HA & for that we have to bring down the SP below 50% each. So I have been running NAR logs & evaluating performance at lun level. We used to see heavy forced flush but now it spikes during the day sometimes & looks much better than before. I have 5 days NAR logs & just used mitrend site to provide automated report on the NAR logs which gives really good performance & capacity output of the array. Now here I have got confused. 1) Firstly theirs no performance problem in our environment 2) In the performance section of the ppt file the NAR output shows some luns as highly utilized lun, then it shows top luns by ABQL, top luns by response time, top luns by forced flushes, top luns by disk crossing, top luns by IOPs.
Now I know top luns by IOPS & diskcrossing can be neglected at the moment because the underlying spindles are able to process the load but I am getting confused by the other outputs like top ABQL, top response time, top forced flushes. Bearing in mind that I am troubleshooting high CPU utilization of the array I can straight migrate the top luns with IOPS by sancopy or Vmware storage Vmotion to the VNX array & this will reduce the amount of load on cx4-960 eventually bringing down the CPU as we have practically seen that earlier. Can anyone recommend a mid way to address this just for the sake if I select some top luns with high IOPS & move it to VNX but if the top ABQL luns are not one of the top IOPS luns which means the dirt still resides on my CX4-960 even though I achieve my target of bringing down the high CPU. Are their any other areas I should look at without just blindly move the top IOPS luns to my VNX
basic details of the array
Flare : 4.30.xx.xx.720
EFD drives count : 9 drives of 100Gb each (used as fast cache)
rest of the drives are 450 GB FC 15k rpm & 600 GB FC 15k RPM
Host environment : ESX 4.0 / windows 2008 (80% of host environment is running ESX4.0 with SQL instances)
I am also attaching the NAR automated ouput provided by mitrend of my array which will clear give more understanding to my query. The slides which I am referring to is from slide no 62 to 84
Thanks for the response. But I am really not looking for a solution to any problem as we don't see any performance issues on the array. But I need some inputs based on the NAR logs output attached to the discussion
For performance issues and to analyze NAR files it is best recommended to open a Service Request and get the files reviewed by the Performance Team for the findings and solutions/recommendation ….
firozg
41 Posts
206
0
Posted June 26th, 2012 05:00
Hi Suman
Thanks for the response. But I am really not looking for a solution to any problem as we don't see any performance issues on the array. But I need some inputs based on the NAR logs output attached to the discussion
regards
Firoz