Unsolved
This post is more than 5 years old
11 Legend
•
20.4K Posts
•
87.4K Points
0
1897
December 23rd, 2009 18:00
What is causing high WP counts ?
Hello guys,
We have a daily BCV job that runs everyday at 5pm. The job typically completes in 30 minutes or so but yesterday after 2 hours it was still running. As we started looking at the frame we noticed one Windows box that was generating crazy WP counts to its 500G LUN (striped meta). I have been staring at these PM graph but can't pinpoint exactly what this box is doing that's generating such high WP numbers. These high WP numbers are observed even before BCV job started.
DMX4-24
Microcode Version (Number) : 5773 (168D0000)
Microcode Patch Level : 138
Cache Size (Mirrored) : 98304 (MB)
# of Available Cache Slots : 922558
# of PermaCache Slots In Use : 26
Max # of System Write Pending Slots : 738767
Max # of DA Write Pending Slots : 369383
Max # of Device Write Pending Slots : 36935
This is 500G striped meta composed of 60 x 8.6g 7+1R5 devices. These devices reside in disk group with 558 x 300G 10K drives so really wide spread.
This is system total WP tracks, when it gets to >50% of 738767 BCV job is hosed
These are WP counts for the 60 members of this meta
Total IOPS these meta members are generating, so 60 x 15 is 900 IOPS. I would think this should be nothing for striped meta spread so wide ?
It's predomenantly write intensive
Not moving a whole lot of data ~36MB/s for read/writes ...that's nothing.
I/O size is pretty big which should be good i would think
I don't understand what could be causing such high WP numbers if IOPS are small, write size is big. What other metrics could give me some hints ?
Thanks a lot


Quincy561
1.3K Posts
0
December 23rd, 2009 18:00
What else is on the disks that are getting the high WP counts? Maybe they are having trouble destaging because of some other competing workload.
Cache partitioning could help isolate the writes, but I think we need to understand what is causing the high WP counts first.
Could you post the BTP file somewhere?
dynamox
11 Legend
•
20.4K Posts
•
87.4K Points
0
December 23rd, 2009 20:00
Quincy56,
i am uploading ttp file to ftp.emc.com/incoming/dynamox
Thank you so much for taking your time to look at it. In the mean time i will look at stats for physical spindles.
dynamox
11 Legend
•
20.4K Posts
•
87.4K Points
0
December 23rd, 2009 21:00
this meta (head 298C) and members are spread over 400 disks. Not sure if there are specific metrics i should look for but looking at %util and %disk busy nothing jumps out
Quincy561
1.3K Posts
0
December 24th, 2009 07:00
So I assume the timeframe you are trying to fix is after 17:00 in the BTP file. If I look at the top 40 devices with the higest WP count, none are in that timeframe.
Is it the batch job you are waiting for or the copy task to complete?
If it is the copy job, and the targets are not true BCVs, I would suggest early activation of the clone targets.
Still not sure what is causing the high write pending counts during that time. Is there by chance adaptive copy RDF in the box that is in WP mode instead of disk mode?
dynamox
11 Legend
•
20.4K Posts
•
87.4K Points
0
December 24th, 2009 09:00
this is the meta where WP start going up around 16:45.
This meta is presented to a windows box where they run some kind of batch job, pulling records from Centera CUA. My priority is to make sure the 5pm BCV job completes on time but their CUA batch job is bringing WP so high that my BCV job just stands still. Their batch can probably be scheduled at different time so it would not affect by BCV job but i am just trying to understand what is it that they are doing that causes such high WP.
My BCV job consists of 45 STD (2-way mirror) and 45 BCV (RAID-5), so there is clone emulation involved but it's never been a problem, incremental sync always completes in 30 minutes or so (270G worth of incremental changes). These 45 STD are running in SRDF Adaptive Copy Disk mode (C.D).
Thank you
Quincy561
1.3K Posts
0
December 29th, 2009 07:00
I can't see a smoking gun that is causing the higher than normal WP counts. I would suggest you open a case, if there isn't already one opened.
John
JayPat1
3 Posts
0
May 19th, 2011 08:00
Hi,
I am seeing the exact same thing on few of my frames. Was anything identified that could bring these wp counts high?
EXACT same scenerio. Batch job, WP count increasing to a point where it hits device max. And the cycle is elongated.
Appreciate your help.
Quincy561
1.3K Posts
0
May 19th, 2011 09:00
In general, if you have high WP counts, you are filling cache faster than the box can destage it.
Solutions can be:
Spread the workload over more drives
Add more DAs and drives
Change the protection
Slow the host workload down
dynamox
11 Legend
•
20.4K Posts
•
87.4K Points
0
May 19th, 2011 17:00
could not pin point exactly what the window box was doing but application people changed the "thread" count in their app and problem went away. I hate situations where i can't explain what happened but BCV cycle was back to normal so everybody is happy.
JayPat1
3 Posts
0
June 9th, 2011 12:00
Thanks Dynamox,
I will try to research on the applications behaviour. My problem still persists. We moved to a different FA' migrated to less used devices. But still persists.
Regards
Quincy561
1.3K Posts
1
June 9th, 2011 13:00
In general, high WP counts are because you are filling cache faster than the disks and DAs can destage them.
If you want to lower the WP counts, add disks, go to faster disks, add DAs, or change the protection to one with less overhead on writes.
JasonBailey
147 Posts
0
July 9th, 2011 18:00
Another path to pursue is to enable Optimizer if you haven't already
dynamox
11 Legend
•
20.4K Posts
•
87.4K Points
0
July 12th, 2011 20:00
not licensed for Optimizer