Unsolved
This post is more than 5 years old
1 Message
0
7704
February 2nd, 2015 02:00
San headquarters enigma...
Hi,
We have a PS4000, loaded with 15K 600Gb SAS drives configured as RAID10. The unit is pretty quiet (in the main - see below) only a couple of Poweredge R710s connected running SQL 2008 active active cluster. We also have a PV6500 that the nodes use for stuff like volumes to perform sql backup to.
I say it's quiet. It is apart from one volume that holds our datawarehouse. This is volume with about 1 TB of SQL database data (log files and temp db are all on other volumes). I'm trying to get the bottom of the latency issues that you can see in this screen grab...
Here is a grab of the disks on the unit...
now here is the thing - that volume is clearly suffering from latency problems - and in the second grab, you can see the disks are busy enough, but I have never, ever seen the disk queue on the volume (the yellow line at the bottom) show anything else other than 0, even with high MB/s IOPS Latency etc. How can this be? Can I trust the san hq data for this volume? I have a ticket open with EQL, and although I'm getting valid prompts to upgrade hit kit, drivers etc, this problem has always pervaded, even when the current drivers / hit kit WERE the latest version! I have done the TCP tweaks for Nagles and TCP delay Ack, but no change.. TCP re-transmits well below anything to worry about.
The guys in support are asking to look at the SQL setup and I have a few whitepapers - all good and very useful, but I can't help thinking there is something screwy with the SAN HQ profile for this volume
Any ideas?

