Unsolved

This post is more than 5 years old

207 Posts

1806

June 9th, 2008 16:00

New to Navisphere Analyzer

Hello,

I spend most of my time working with DMX storage and I am familiar with ECC Performance manager for symmetrix.

Today I am looking at some issues with AIX hosts that are connected to a cx380 with 8GB cache. These hosts are used for Data Warehouse, SAS, Filenet and a large reporting application. The Filenet host experiences slow response time every once in a while. I'm not sure yet whether this is a host, application, database or storage issue.

In Navisphere I turned on logging and now I can see some statistics with Analyzer. In the "Performance Overview - Bandwidth" screen, I can see that the total MB/s is around 300 MBs most of the time today. The IOPS are around 4000. Comparing this to my symmetrix arrays it is pretty high - my symm arrays are usually around 50Mbs with 2,500 IOPS. Is 300MBs high for a cx-380? I'm not familiar with performance characteristics for a cx.

Also, I have noticed in Analyzer "Performance Summary" that three LUNs on the SAS host are almost always 100% utilized during the day. They sit on the same RAID 5 Group with several other LUNS. It is made up of 4 disks. I'm guessing this RAID Group and the LUNs would benifit if we added more disks. Maybe add an additional 4 disks to this RAID group. Any thoughts?

Also, I am thinking cache may get flooded when these hosts are pounding the cx. Eight GB of cache doesn't seem like much - my symms have 64GB cache. Is there some statistic that shows cache hit ratio and other info on cache usage?

207 Posts

June 9th, 2008 16:00

More info as I continue to look with Analyzer - "Dirty Page %" is averaging about 75%. This seems high to me??

207 Posts

June 9th, 2008 16:00

As I look further I can see that many of the Data Warehouse LUNs are up around 90% utilization.

133 Posts

June 10th, 2008 02:00

Hi BradRondeau,

Last things first ;-)

The %Dirty Pages ideally fluctuates between the low and high water marks. The default values are 60 and 80%. Forced flushing, emphasises clearing cache over front end activity, is what you need to keep an eye out for which occurs at 100%. Forced flushing in a production environment per LUN should be < 10/second

By the sound of things the CLARiiON has got its work cut out! But before jumping to conclusions whether you should add more spindles to your RGs or migrate the LUNS to different RGs/RAID levels you should do a full performance evaluation of the array starting with the SPs followed by the LUNS and disks.

Some of the things you should verity are:

SP
- Utilization, < 70% and evenly between SPA and SPB
- Cache settings, unless instructed by analyzer you should max the write cache
- Dirty Pages


LUN
- Utilization
- Cache, enabled?
- Trespassing
- Forced flushes
- IO size, hitting the cache or larger then write aside
- Read throughput & I/O size
- Average busy queue larger than queue depth > bursty app > more spindles

Disks
- Utilization
- Disk IOPS:
 Over 140 for small I/O (FC)?
- Disk MB/s:
 Over 10 for large I/O (8 MB/s ATA)?
- Disk Service Time
- Disk Queue Lengths and Disk Response Time
- Disk Average Seek Distance, <1/4 is good & >1/2 needs further investigation
- Coalescing, disk reads/writes larger than LUN reads/writes?

In addition have you got any layered apps in the mix: snaps, clones, mirrors¿?

If you got the chance book yourself on the excellent CLARiiON performance workshop as well!

I hope this helps
Carl

6 Operator

 • 

4.5K Posts

June 10th, 2008 10:00

In addition to the suggestions from calli, I would also recommend that you read the following two documents:

EMC CLARiiON Best Practices for Fibre Channel Storage: FLARE Release 26 Firmware Update - Best Practices Planning

http://powerlink.emc.com/km/live1/en_US/Offering_Technical/White_Paper/H2358_clariion_best_prac_fibre_chnl_wp_ldv.pdf

EMC CLARiiON Fibre Channel Storage Fundamentals - Technology Concepts and Business Considerations

http://powerlink.emc.com/km/live1/en_US/Offering_Technical/White_Paper/H1049_emc_clariion_fibre_channel_storage_fundamentals_ldv.pdf

In order to determine what you need, you need to use the archives that are created when using the Analyzer - there's an option called "Periodic Archiving" turn this on and there will be an archive created very n hours - n = the number of hours based on the archive interval. If the archive interval is set to 120 seconds, then a new archive will be created about every 5 hours. If you set the archive interval to 60 seconds, the archives will hold about 2.5 hours of data.

Go to Tools/Analyzer/Customize and on the General tab check the Advanced box. This will show you more metrics then the default view.

Once you have the archives you can retrieve and open them using analyzer. Go to Tools/Analyzer/Archive/Retrieve to download the archives from the array to your workstation. Then use Tool/Analyzer/Archive/Open and point to the archive to open.

The on-line Help in Navisphere is very good - it explains what the different metrics mean and how to location bottlenecks.

With disks you want to look at the total IOPS for each disk - the Best Practices has a section called "Sizing the Storage requirements" around page 44 that lists the different disks and the IOPS for best performance. In general, if you have a 10K FC disk it should be able to handle about 120 IOPS and still provide good performance (depending on IO size, etc.). If you see that the IO to each disk is exceeding this, you may need to add more disks to the Raid Group. All this depends on a lot of factors.

regards,

glen kelley

207 Posts

June 11th, 2008 16:00

Thanks for the info. I will look at the links provided for additional info. I did turn on the advanced flag. Retreiving the .nar files was strange when I was coming into Navi via my domain. I had to come in to Navi using the IP of the specific array in order to get at the .nar files. I will start looking at them tomorrow.

It takes a while for Analyzer to get all the info because I have 250 Luns, 360 disks, many RAID groups etc. Sometimes Analyzer seems to hang or freeze. I'm more comfortable with ECC Performance Monitor so I'm hoping I can get at my CX data there once we upgrade to ECC 6.

Currently the SPs are both below 50% usage. I wish that I would have checked this on Monday when the array was much busier. On Monday the MBs was around 270. Would that be considered high usage for a cx-380? We have another big application to add and I'm not sure whether we are running out of horse power on this cx.

6 Operator

 • 

4.5K Posts

June 12th, 2008 11:00

270 MB/s is not too high for a CX3-80, but it all depends on the environment - IO size, number of paths from the host to the array, number of threads, etc. The Best Practices guide should be able to help in determining the upper limits on load.

glen

207 Posts

June 12th, 2008 16:00

Good info - I need will look into cx training. Also, I am not running any layered apps.

I am running into issues with Analyzer running extremely slow and locking up. Is there another way to view the data more efficiently.

Also, my Write Cache Flushes are around 100 per second sometimes reaching 250. My Average Busy Queue Length is around 10 sometimes as high as 25. Are these issues?

207 Posts

June 13th, 2008 09:00

On ECC Performance manager for the dmx their is a counter called "Cache Hit Ratio %". Is there a similar counter in Analyzer? I have not been able to find it.

6 Operator

 • 

4.5K Posts

June 13th, 2008 12:00

Are you looking at the Real-Time Analyzer or the Archive (NAR) files? NAR files are the best and impose the least overhead on the array.

If you're seeing 100% Dirty Pages, you have something serious going on - have you checked the SP Event log to see if you're getting excessive trespass events or seeing something like "Background Verify Started" / "Background Verify Aborted" messages? The Background Verify message are specifically for Veritas DMP - if you see these messages, shot your Veritas servers to stop.

The other issue with 100% Dirty Pages may be a backup from Fibre LUN's to ATA LUNs - the Write to the ATA disks are filling the Write Cache.

Go to Tools/Analyzer/Customize - on the General tab, check the Advance settings box.

There are a bunch of cache counters - Read Cache Hit Ratio, Write Cache Hit Ratio, Cache Misses, etc. - you need to be looking at a LUN and the LUN selected (highlighted) in the upper left box - the one with all the LUN's listed. The counters in the lower left box of Analyzer will change depending on where the focus is - SP or LUN or Disk - each has similar and different counters.

glen

6 Operator

 • 

4.5K Posts

June 13th, 2008 12:00

Write cache flushes occur all the time - when the Write cache hits the High Watermark, the cache flushes to disk. There is also a idle write cache flush that periodically flushes the cache. You should check to see if you're hitting 99% Dirty Pages - a few are OK, but more than that are a problem - see my other note above.

glen
No Events found!

Top