Unsolved
This post is more than 5 years old
207 Posts
0
1806
June 9th, 2008 16:00
New to Navisphere Analyzer
Hello,
I spend most of my time working with DMX storage and I am familiar with ECC Performance manager for symmetrix.
Today I am looking at some issues with AIX hosts that are connected to a cx380 with 8GB cache. These hosts are used for Data Warehouse, SAS, Filenet and a large reporting application. The Filenet host experiences slow response time every once in a while. I'm not sure yet whether this is a host, application, database or storage issue.
In Navisphere I turned on logging and now I can see some statistics with Analyzer. In the "Performance Overview - Bandwidth" screen, I can see that the total MB/s is around 300 MBs most of the time today. The IOPS are around 4000. Comparing this to my symmetrix arrays it is pretty high - my symm arrays are usually around 50Mbs with 2,500 IOPS. Is 300MBs high for a cx-380? I'm not familiar with performance characteristics for a cx.
Also, I have noticed in Analyzer "Performance Summary" that three LUNs on the SAS host are almost always 100% utilized during the day. They sit on the same RAID 5 Group with several other LUNS. It is made up of 4 disks. I'm guessing this RAID Group and the LUNs would benifit if we added more disks. Maybe add an additional 4 disks to this RAID group. Any thoughts?
Also, I am thinking cache may get flooded when these hosts are pounding the cx. Eight GB of cache doesn't seem like much - my symms have 64GB cache. Is there some statistic that shows cache hit ratio and other info on cache usage?
I spend most of my time working with DMX storage and I am familiar with ECC Performance manager for symmetrix.
Today I am looking at some issues with AIX hosts that are connected to a cx380 with 8GB cache. These hosts are used for Data Warehouse, SAS, Filenet and a large reporting application. The Filenet host experiences slow response time every once in a while. I'm not sure yet whether this is a host, application, database or storage issue.
In Navisphere I turned on logging and now I can see some statistics with Analyzer. In the "Performance Overview - Bandwidth" screen, I can see that the total MB/s is around 300 MBs most of the time today. The IOPS are around 4000. Comparing this to my symmetrix arrays it is pretty high - my symm arrays are usually around 50Mbs with 2,500 IOPS. Is 300MBs high for a cx-380? I'm not familiar with performance characteristics for a cx.
Also, I have noticed in Analyzer "Performance Summary" that three LUNs on the SAS host are almost always 100% utilized during the day. They sit on the same RAID 5 Group with several other LUNS. It is made up of 4 disks. I'm guessing this RAID Group and the LUNs would benifit if we added more disks. Maybe add an additional 4 disks to this RAID group. Any thoughts?
Also, I am thinking cache may get flooded when these hosts are pounding the cx. Eight GB of cache doesn't seem like much - my symms have 64GB cache. Is there some statistic that shows cache hit ratio and other info on cache usage?
No Events found!


brad12341
207 Posts
0
June 9th, 2008 16:00
brad12341
207 Posts
0
June 9th, 2008 16:00
calle2
133 Posts
0
June 10th, 2008 02:00
Last things first
The %Dirty Pages ideally fluctuates between the low and high water marks. The default values are 60 and 80%. Forced flushing, emphasises clearing cache over front end activity, is what you need to keep an eye out for which occurs at 100%. Forced flushing in a production environment per LUN should be < 10/second
By the sound of things the CLARiiON has got its work cut out! But before jumping to conclusions whether you should add more spindles to your RGs or migrate the LUNS to different RGs/RAID levels you should do a full performance evaluation of the array starting with the SPs followed by the LUNS and disks.
Some of the things you should verity are:
SP
- Utilization, < 70% and evenly between SPA and SPB
- Cache settings, unless instructed by analyzer you should max the write cache
- Dirty Pages
LUN
- Utilization
- Cache, enabled?
- Trespassing
- Forced flushes
- IO size, hitting the cache or larger then write aside
- Read throughput & I/O size
- Average busy queue larger than queue depth > bursty app > more spindles
Disks
- Utilization
- Disk IOPS:
Over 140 for small I/O (FC)?
- Disk MB/s:
Over 10 for large I/O (8 MB/s ATA)?
- Disk Service Time
- Disk Queue Lengths and Disk Response Time
- Disk Average Seek Distance, <1/4 is good & >1/2 needs further investigation
- Coalescing, disk reads/writes larger than LUN reads/writes?
In addition have you got any layered apps in the mix: snaps, clones, mirrors¿?
If you got the chance book yourself on the excellent CLARiiON performance workshop as well!
I hope this helps
Carl
kelleg
6 Operator
•
4.5K Posts
1
June 10th, 2008 10:00
EMC CLARiiON Best Practices for Fibre Channel Storage: FLARE Release 26 Firmware Update - Best Practices Planning
http://powerlink.emc.com/km/live1/en_US/Offering_Technical/White_Paper/H2358_clariion_best_prac_fibre_chnl_wp_ldv.pdf
EMC CLARiiON Fibre Channel Storage Fundamentals - Technology Concepts and Business Considerations
http://powerlink.emc.com/km/live1/en_US/Offering_Technical/White_Paper/H1049_emc_clariion_fibre_channel_storage_fundamentals_ldv.pdf
In order to determine what you need, you need to use the archives that are created when using the Analyzer - there's an option called "Periodic Archiving" turn this on and there will be an archive created very n hours - n = the number of hours based on the archive interval. If the archive interval is set to 120 seconds, then a new archive will be created about every 5 hours. If you set the archive interval to 60 seconds, the archives will hold about 2.5 hours of data.
Go to Tools/Analyzer/Customize and on the General tab check the Advanced box. This will show you more metrics then the default view.
Once you have the archives you can retrieve and open them using analyzer. Go to Tools/Analyzer/Archive/Retrieve to download the archives from the array to your workstation. Then use Tool/Analyzer/Archive/Open and point to the archive to open.
The on-line Help in Navisphere is very good - it explains what the different metrics mean and how to location bottlenecks.
With disks you want to look at the total IOPS for each disk - the Best Practices has a section called "Sizing the Storage requirements" around page 44 that lists the different disks and the IOPS for best performance. In general, if you have a 10K FC disk it should be able to handle about 120 IOPS and still provide good performance (depending on IO size, etc.). If you see that the IO to each disk is exceeding this, you may need to add more disks to the Raid Group. All this depends on a lot of factors.
regards,
glen kelley
brad12341
207 Posts
0
June 11th, 2008 16:00
It takes a while for Analyzer to get all the info because I have 250 Luns, 360 disks, many RAID groups etc. Sometimes Analyzer seems to hang or freeze. I'm more comfortable with ECC Performance Monitor so I'm hoping I can get at my CX data there once we upgrade to ECC 6.
Currently the SPs are both below 50% usage. I wish that I would have checked this on Monday when the array was much busier. On Monday the MBs was around 270. Would that be considered high usage for a cx-380? We have another big application to add and I'm not sure whether we are running out of horse power on this cx.
kelleg
6 Operator
•
4.5K Posts
0
June 12th, 2008 11:00
glen
brad12341
207 Posts
0
June 12th, 2008 16:00
I am running into issues with Analyzer running extremely slow and locking up. Is there another way to view the data more efficiently.
Also, my Write Cache Flushes are around 100 per second sometimes reaching 250. My Average Busy Queue Length is around 10 sometimes as high as 25. Are these issues?
brad12341
207 Posts
0
June 13th, 2008 09:00
kelleg
6 Operator
•
4.5K Posts
0
June 13th, 2008 12:00
If you're seeing 100% Dirty Pages, you have something serious going on - have you checked the SP Event log to see if you're getting excessive trespass events or seeing something like "Background Verify Started" / "Background Verify Aborted" messages? The Background Verify message are specifically for Veritas DMP - if you see these messages, shot your Veritas servers to stop.
The other issue with 100% Dirty Pages may be a backup from Fibre LUN's to ATA LUNs - the Write to the ATA disks are filling the Write Cache.
Go to Tools/Analyzer/Customize - on the General tab, check the Advance settings box.
There are a bunch of cache counters - Read Cache Hit Ratio, Write Cache Hit Ratio, Cache Misses, etc. - you need to be looking at a LUN and the LUN selected (highlighted) in the upper left box - the one with all the LUN's listed. The counters in the lower left box of Analyzer will change depending on where the focus is - SP or LUN or Disk - each has similar and different counters.
glen
kelleg
6 Operator
•
4.5K Posts
0
June 13th, 2008 12:00
glen