tthomas1

updated

16 years ago

T

tthomas1

10 Posts

0

7815

September 10th, 2010 13:00

Understanding LUN Performance

Could someone help me understand how to interpret LUN performance in Navi Analyzer?  I have been doing some reading on past discussions but I'm looking for a clearer answer if there is one.  My situation begins with a SQL 2005 DB cluster.  There have been about three instances over the last few months where the cluster has failed from the active node to the passive node.  Each time the storage performance comes into question sighting disk response time per the software used to monitor the cluster (Idera).  According to the Analyzer data I don't see any major problems.  The last time the failover occurred I did notice that the LUN containing the DB data (RAID 1/0, 4+4, dedicated) appeared to experience a high response time (64 ms) around the time of the failover.  The LUN utilization and queue length (25% and 4.7 respectively) are rather low, at least as far as my knowledge is concerned.  Forced flushes for that LUN are practically nothing during that same time period.  The SPs' performance during that time period are as follows: Utilization 14 and 28%; Queue length 2.5; Resp time 1.75 ms; Dirty pages SPA 71%, SPB 89% (watermarks 80/60).

Now that I've said all that, I've read that LUN response time can or is likely to be falsely reported during times when the utilization and total throughput for that LUN are low and that you should only be concerned when all three of those metrics are high at the same time.  Is this true?  If so, is LUN response time a reliable metric to monitor?  Some of these metrics seem to be a challenge to understandm, at least for me.

I apologize for the lengthy explanation but I wanted to be sure that whomever took this one on had the necessary details.  Any help will be appreciated.

Thanks,

Tim