Hi -- I'm looking for a way to summarize daily values for an array, going through the .nar files converted to .csv. I understand the TotalIOPS for the SPA/B to be values for when the array is polled, so that wouldn't work. The doc seems to say that the TotalIOPS for the front-end ports on the SPs are total values. So, if I add those values over a 24 period, would that give me what I'm looking for? The number seems low.
Hi. Analyzer samples every poll period - that can be 60 seconds to 3600 seconds (1 hour). Let's say you have 120 second sampling.
So you collect nar files and most metrics for all objects are averages for the sample period - you dump the stats (maybe using naviseccli analyzer -archivedump) and you get your csv format files.
As you point out, the label is SP Total Throughput, but it does state IO/s, so it is actually an average throughput for the sample period.
So if your sample period by the data logging is 120 seconds, then you want to take each value, and multiply by the 120 to get total I/O for that period e.g. if you see the SP reporting 17,086 IO/s (example I am looking at), that one SP would have processed on average 17,086 * 120 = 2,050,320 I/O in the last 2 minutes. You woudl do this math for each sample then aggregate all values to cover the 24 hour period if that is hat you are looking for - if this was a consistent benchmark, just as an example, my 17,086 is almost constant right now so for 24 hours, that value would equate to nearly 1.5 billion I/O in one day - that's probably more like the number you are expecting, and that was just for one SP.
Just a few aditional notes: SP stats won't reflect what is going in at the backend as far as disk IOPs as you won't be factoring in RAID overheads, internal operations for things like dedupe, compression, migrations, etc. These operations are normlly available at the disk level but are difficult to isolate from normal host I/O so probably best to avoid unless you really want to be digging, and that can really be a challenge
The values are an average for the sample period - as the actual value at any point in time can fluctuate and most likley will in a variable workload environment, but as we're looking at multiplying the average for the sample period, and looking at totals here, you should be in good shape. I only point that out as you could see a value of 10,000 but at times the SP might go much higher and have periods that are much lower i.e. average comes in lower, but we don't report the range of IO as that isn't a viable option and why we have the variable sample period - this should not affect what you want to do though.
For determining if we have variation, you have to look at the LUN stats and compare average queue and average busy queue i.e. if the average busy queue is higher, that could imply we had more work being done within the sample period that the average doesn't reveal the peak for - it's one of those things that just gives us a clue the LUN is busier or sometimes idle within the period time but as we're not tracking every I/O, we can't report the peak (to track every I/O that would be very intensive work for non-user I/O).
Hi Susanne. Well, it's kind of misleading wording - you can't have an IO/s value for a point in time as it is what it says it is: I/O per second, so it has to represent an average throughput for some time period. There is no such thing as an instantaneous performance value.
So what it is telling you, is that at the time of the poll, the I/O per second does actually represent the average for the poll period but the average right now. One could say the same value would be the average for anywhere within that poll period, as long as the value considered the time up to and including the end of the poll period.
It has always been a little grey as to pople's interpretation as it's rarely explicitly stated it's the average for that period. It's just interesting they chose, within this document, to state for some metrics "at the time it is polled" versus other metrics where it's not mentioned.
I hope that clarifies that (sorry for the wording and/or reference inconsistency in that document)
Now, as you mentiond, you could aggregate the per port I/O per socond values and do the math, but you should get the same value if you take the SP total I/O per second value and multiply by the poll time - I did that yesterday to double check and they matched. Obvioulsy you can try that to give yourself confidence to use that single value rather than dumping the port stats and doing more work, albeit logical number manipulation you are probably automating somehow
Jyothi_P_Bharat
1 Rookie
•
317 Posts
316
1
Posted July 3rd, 2014 02:00
Hi,
I think the below thread may help you,but not sure:
https://community.emc.com/message/817865#817865
Thanks
Jyothi