
UNSOLVED
VNX7500 Array - Citrix PVS 6.5 sluggish performance
Hi
We have got a VNX Array 7500 with 900 GB of Fast cache / 8 GB of write cache on each SP. with 600 drives running. 8 x 30 drives pools of RAID 5 and ESX host 4.0 connected via Cisco MDS 9500 series. Citrix PVS is deployed during the initial phase to support approx 600-700 users. All the Luns on the array are thick luns assigned to ESX host with a standard size of 500 GB each. Total of 90 Xenapp servers for users who connect through and a cluster PVS server. User experience is not good and the problem is intermittent without any pattern. Some users face slowness while logging in but some dont. some have issues in accessing their files but some don't. Theirs no definitive problem but a in-consistent and random issues which all the users compliain that is sluggish performance. I have run NAR files to two days and this is what I found
total of 31 luns are in question here and all RAID 5 thick luns. Total IOPS of these luns are 12000 approx and R:W ratio is 4000 : 8000 IOPS i.e. 1:2. SP utilization is slightly lower than 40 both the SP. memory Cache is ranging between 60 and 90 as oppose to (LWM/HWM which is set to 60-80) At times it does reaches 100 but once or twice a day for a couple of seconds. Pool lun response times are ranging from 5ms to 15ms except few spikes which goes upto 35ms. ABQL for all the luns are below 20 majorly between 5 and 10. SP response time is 1-2 ms (both SP). VNX array is running flare 32p11. Analyzer doesnt provide forced flushes for pool luns but found it is good if pool luns are below 20 ms.
couple of forced luns are found but those are RG luns and spikes are not throughout just a couple of times during the day.
Some drives in the pool has ABQL over 1 but again those are spiky.
Recommendations are made as below
1) migrate all the Xenapp VMs to RAID 10 as the Writes are 66% and reads 33%
2) Stop mirror view replication of these luns
Lun Numbers in questions are (Lun numbers 10 to 33 and 250 to 251)
Questions I have is:
The stats I stated above doesnt really shows any major problem on the array but the number of users affected is huge. Not sure if this really is a SAN issue or their are other factors. Dont know if moving the VMs from RAID 5 to RAID 10 will really improve user experience.
I have attached the NAR files for 31st Jan'13 and 1st Feb'13 to this forum. Will be oblige if anyone can really have a look and let me know if the recommendations will help
Working business hours are only from 08:00 hrs to 18:00 hrs GMT
2 Attachments
CKM00114800241_SPA_2013-01-31_20-19-08-GMT_P00-00.nar
CKM00114800241_SPA_2013-01-31_20-19-08-GMT_P00-00.nar
01_Feb_13_merged.nar
01_Feb_13_merged.nar
Responses (0)
Solutions (0)
