
Solved!
Go to SolutionApplication slowness issue
Hi All
I am facing issue in analyzing the problem with one of the 1000 VMs on our CX4-960. We got ESX4.0 with 1000+ VMs in multiple clusters. Linux RHEL 4.0 on an ESX 4.0 connected to CX4-960 4.30.xx.xx.525, This VM is mounted on a 6 different luns with different raidgroups on the array. The VM has an Oracle DB running on it. The problem statement is they see high IO wait only during 7:00 to 08:00 am GMT & thats the time when 2 different activity happens (1) Users login in (2) Oracle schedule jobs are run. The rest of the day it is running perfectly fine with no issues. Server team claims CPU/Memory of the server to be under control & only bottleneck they see is IOwait. So I ran NAR files for couple of days and this is what I see.
Out of the 6 luns their are 2 luns of RAID 10. One of the lun is on a 4+4 RG where as other one is a metalun spread on a two different RGs of (4+4 i.e. 16 spindles) I have analyzed all the 6 luns and found only one 350 GB luns which is a metalun residing on a 2 x 8 spindles RAID 10 RG
This is a striped metalun of 175 GB each carved on a (4+4) 2 different RG R1_0 (FC 15krpm drives)
Below is the output i see on the NAR files
Lun utilization during the slowness period is 100%
Lun total throughput is 160 IOPS
r/w ratios is 80:80 iops
though read bandwidth is hardly an mbps but writes are big chunks upto 8-10 mbps with filesize over 128kbps
Lun doesn't show SP forced flush not even once
SP utilization during this period is hardly 45% (both SPs with sp cache being within 60-80 watermark) no flush occuring
Lun queue length is between 8 to 14 same as ABQL
Lun response time seems to be around 300 ms but i dont think this is any problem as its been processed by cache
I drilled down to the Raidgroups/disks and found it to be very optimal and underutilized
Disk utilization of both the RGs is under 10%
Disk queue length of both RGs is 0.50 to 0.80
Disk response tims is ranging from 15 ms to 69 ms (this i feel is an issue)
each disks in both the RGs are transfering @ the rate of 1-2mbps
IOPS of each drive is not above 40 IOPS
write files size on each drive seems higher than 96 kbps.
read & write bandwidth both are below 1 mbps
ABQL is between 8 to 12
I verified all the other luns in the two RGs are very quite during the problem period basically nails down to this lun & its component lun in the other RGs
How would I judge if this lun really has issues in getting IOs processed. Though my Disk respone time is higher but i still dont see a single spike on my SP Cache. Write file size are bigger and so the bandwidth too is higher on the lun but the same is low in my drives.
Is their any other parameter which can give me a clue here? its been 3 days since the problem is appearing & its hitting hard in my nerves.
I am not knowledgeable on Linux stuff and the server owner loaded me with the IO outputs when asked and was only able to drill down to the devices SDC/SDC1. I have attached the output in the text format to this discussion.
This lun is in mirrorview/A relationship
any advice at the earliest to this will be highly appreciated
thanks
Firoz
3 Attachments
io.out26112012
io.out26112012
WORKLOAD REPOSITORY report for UKPRHRDB02.docx
WORKLOAD REPOSITORY report for UKPRHRDB02.docx
CKM00094500211_SPA_2012-11-28_10-46-43-GMT_P00-00.nar
CKM00094500211_SPA_2012-11-28_10-46-43-GMT_P00-00.nar
Responses (0)
Solutions (0)
