Unsolved
This post is more than 5 years old
15 Posts
0
1029
March 29th, 2007 10:00
CX3-80 strange performance metrics and disk access issues
I have a new CX3-80 installed which is replacing my CX500. I have moved over one 400GB parition using a backup/restore method and have had a number of strange problems since. I don't know if the NaviAnalyzer stuff is much different on the 80, but here is what I am seeing:
Collected Data = SPB, LUN1100, watching Service Time and Write Cache Hit Ratio. Approximately every 4 minutes, all of these metrics go from their average value down to 0. Sometimes it goes a little longer than 4 minutes, but generally it almost looks like some type of flush event.
If I monitor a disk in LUN1100, the service time does not drop to 0. During that same time, the SP Queue length stays normal (.3), but the disk queue length drops to 0 with the service time on the LUN.
The problem that we are having is that a Windows2003E host reports "Not enough server storage is available to process the command" during large file transfer operations (100+mb generally). After it happens once for the day, we generally need to fail the resource over to a new server (MS Clustering).
We have identical traffic on our CX500. We even built a new cluster node using the same (older) server hardware vs. the new stuff we put in to use with the new SAN. Same drivers, same HBA's, same PP, etc.
The new SAN is connected to a new Cisco MDS, while the CX500 is using older 32 port Brocades.
Any ideas? Has anyone seen any performance metrics like that and if so, is that considered normal?
Collected Data = SPB, LUN1100, watching Service Time and Write Cache Hit Ratio. Approximately every 4 minutes, all of these metrics go from their average value down to 0. Sometimes it goes a little longer than 4 minutes, but generally it almost looks like some type of flush event.
If I monitor a disk in LUN1100, the service time does not drop to 0. During that same time, the SP Queue length stays normal (.3), but the disk queue length drops to 0 with the service time on the LUN.
The problem that we are having is that a Windows2003E host reports "Not enough server storage is available to process the command" during large file transfer operations (100+mb generally). After it happens once for the day, we generally need to fail the resource over to a new server (MS Clustering).
We have identical traffic on our CX500. We even built a new cluster node using the same (older) server hardware vs. the new stuff we put in to use with the new SAN. Same drivers, same HBA's, same PP, etc.
The new SAN is connected to a new Cisco MDS, while the CX500 is using older 32 port Brocades.
Any ideas? Has anyone seen any performance metrics like that and if so, is that considered normal?
No Events found!


tim0871
15 Posts
0
March 29th, 2007 10:00
I was hoping that maybe someone else would either have come across a similar issue or would have a good idea what else we can look at while we wait. My users get a little frustrated when they keep loosing access to their file stores.
Allen Ward
6 Operator
•
2.1K Posts
0
March 29th, 2007 10:00
Are you still running FLARE 22 on your CX3 or has it been upgraded to FLARE24?
I haven't seen this specifically, but the first thing I would look at if I was seeing this kind of activity would be a possible trespassing/failover fault. Have you checked to see if this LUN is maybe trespassing between SPs? Just off the top of my head, there are a few things that could cause this: Failing port on the switch, Failing port on HBA (depending on how you zone), some kind of software conflict with PowerPath on the host (unlikely though to cause this), or a bad FC cable somewhere in the SAN.
If you have checked the trespassing issue and it is still happening I would go straight to EMC Support with SP Collects from the array and EMCReports generated from the host(s).
tim0871
15 Posts
0
March 29th, 2007 10:00
tim0871
15 Posts
0
March 29th, 2007 11:00
Allen Ward
6 Operator
•
2.1K Posts
0
March 29th, 2007 11:00
I apologize for asking you the same stuff you have already tried, but I haven't seen posts from you before on here so I have no idea of your level of experience. You can never assume anything
Hopefully someone else has something for you, or Engineering gets back to you sooner than later.
tim0871
15 Posts
0
March 29th, 2007 11:00
tim0871
15 Posts
0
March 29th, 2007 11:00
Allen Ward
6 Operator
•
2.1K Posts
0
March 29th, 2007 11:00
I'm not familiar with the Cisco switches, but I'm assuming they have some kind of error tracking on them. Do they show any strange activity on the ports involved when this occurs?
an_hidden_KB
87 Posts
0
March 30th, 2007 02:00
tim0871
15 Posts
0
March 30th, 2007 05:00
You did make me double-chedk my RAID5 groups though. They are all 4+1 and not yet in use.
EKellerman
76 Posts
0
March 30th, 2007 06:00
comprised of 2 sets of 12 drive RAID Groups. The
RAID groups were created in engineering mode so the
mirrored pairs are across different buses.
I'm not fimilar with this....you can actually define where the mirrored pairs are laid out? We don't use 1/0 too often, but this could be useful when we do.
Can you explain this further?
tim0871
15 Posts
0
March 30th, 2007 06:00
naviseccli -h 10.100.10.180 -User [xx] -Password [xx] -Scope 0 createrg 102 0_2_0 1_2_0 0_2_1 1_2_1 0_2_2 1_2_2 0_2_3 1_2_3 0_2_4 1_2_4 0_2_5 1_2_5
tim0871
15 Posts
0
March 30th, 2007 06:00