Unsolved

This post is more than 5 years old

2473

July 9th, 2009 08:00

WP Track utilization and timefinder clones.

I've got two times during the day when the WP track utilization hits 50% for over an hour on my DMX3. I'm concerned about the 50% threshold because at that point my clones stop copying and I'm told the frame makes a system wide priority change to focus on de staging those tracks a little more. I looked at what devices are using the most tracks at that time and it looks like I've got two culprits.

One is a little 46GB oracle database. The four 16.86GB devices assigned to them were using a total of 290,000 WP at peak time. I tried to fix it by throwing more spindles at them and restriped the database over 8 devices. Now during peak time the devices are using 405,000 WP. I increased the throughput the database was doing by 20% so I'm getting more work done but now that little 46GB database is using 25% of my WP tracks on an array with 200TB of usable disk. Is that normal? Any suggestions on what to do?

The other is some timefinder clone targets I use for backup I've been told I can use symqos to lower the priority on the targets but after looking at the manpage and listing the current copy pace I don't really know what I'm looking at. Anyone familiar with symqos that can give me an idea of what I'm looking at / what I should look at doing?

Thanks,
Christian

July 9th, 2009 08:00

I like the idea of the cache partiton and throw all of my clone targets in it.
I've got 128? GB of cache and clone ~ 15TB a day. Do you think 16GB partition would be a good place to start?

1.3K Posts

July 9th, 2009 08:00

There are two things you can do to reduce the WP count on the clone targets.

The first thing is apply a copy pace to the copy process. You must apply the QoS value on the source volumes of the copy. You will need to play with various values to see what reduces the WP count without impacting the copy rate too much. I would try starting with a value of 4, and go up or down from there. The parameter can be changed on the fly at will.

The other thing that can be done is to put the clone target devices into a small cache partition to limit the WP counts. This is actually a preferred solution because it generally doesn't slow the copy at all, and greatly reduces the amount of WP tracks that can build up in the system.

1.3K Posts

July 9th, 2009 09:00

I would start with a 10% cache partition. Set it up with a minimum value of zero. Target = 10%, max = 10%. Leave the default WP value of 80%. This will limit the clone targets to 8% of total cache. The time value could be set to something like 300 seconds or so. This parameter really would depend on what you are doing with the clone targets.

Also don't try to set this up while the clone targets are already at 50%. Wait to put them in a partition when the WP count is low, or zero on them before putting them in a partition.

1.3K Posts

July 9th, 2009 09:00

One other note. Cache partitioning requires a separate license, and clone pace does not.

1.3K Posts

July 9th, 2009 12:00

From your post, it also seems you may have some buildup of WP tracks from the workload and not the copy. While the WP count is below 50%, the back end treats the destage as a low priority. If there are enough read misses to keep the DAs busy, this might be normal. When you reach 50%, destage and read miss become equal priorities, so destage becomes more aggressive. If you are frequently at the 50% because of your normal workload, you probably should be looking into why, because it probably won't take a lot more workload to push over into system WP limits.

1.3K Posts

July 9th, 2009 12:00

In DMX3 and DMX4 each device can consume up to 5% of writable cache. So it is easy to see why just a few devices can get you to the 50% threshold without too much effort. With clones this happens more quickly if your source devices are much faster at reading than your target devices are at destaging. Maybe your source devices are on more disks, or a slower protection type for destage?

You might find a QoS pace setting that won't slow the copy much but keep your WPs under control. The QoS slows the reads, not the writes, so if you can keep the read rate just below what the destage rate is, the system should have a much lower WP count.

July 9th, 2009 12:00

As soon as we activate the clone pair we start backing up the target. Is that what you are talking about by the workload?

July 9th, 2009 12:00

Thanks for the info, I'll se what I can do with the options I have.

Anyone have any thoughts on my little database that is using so much wp?

1.3K Posts

July 9th, 2009 12:00

No, I was more talking about the souce workload without clone active, if you are worried about WP tracks while your clones are not active.

I think we already covered what can be done about WP created by the copy process.

130 Posts

July 10th, 2009 10:00

Checking my system after following this thread.I could see WP tracks behavior where would i find the max limit or matter of fact how would i say it reached 50%?

1.3K Posts

July 10th, 2009 11:00

You can see WP counts with Performance Manager, STP or symstat commands.

You can get your system WP maximum slot count with symcft list -v

July 13th, 2009 06:00

That was the other half of my question at the beginning. I've got that little database that is using up 25% of the WP tracks during the same time period. I'm not sure what to do to bring that down. I'm getting a quote for DCP as we are not licensed. Any other thoughts on how to lower the WP utilization for those 8 devices?

July 13th, 2009 06:00

In performance manager you can map both the WP utilization and Max on the same graph.

1.3K Posts

July 13th, 2009 06:00

25 percent WP isn't anything to be worried about. You want your writes to cache to be "fast" and not to have to wait for destage. If you force them into a smaller partition, chances are you will start hitting partition limits and the host will then have to wait for a destage to be able to write to cache.

Just to be clear, you are seeing this 25% usage when there are no clones active for these devices, correct?

And if you really want to reduce the WP count for these devices, the best way would be to stripe them over more disks and DAs.

July 13th, 2009 08:00

This is while the symm is at 50%. I've got 8 devices striped into a logical volume that are using up ~405K WP which equates to about 25% of the WP tracks for the whole symm. It used to be 4 devices which were using 290K WP tracks, I restriped it over 8 devices and it increased the WP count instead of lowering it. I've already got 135GB usable allocated to a 46GB database. Considering the disk we clone to for backups that is 270GB of disk for a 46GB database. I'm trying to figure out a way to get the overall system usage down without burning up a ton more disk.
No Events found!

Top