Unsolved

This post is more than 5 years old

40 Posts

2559

May 18th, 2006 21:00

SCSI bus resets driving me crazy

Hello all knowledgeable NetWorker Admins. I have been experiencing SCSI bus resets for the past 9 days. My environment is very unstable and I cannot trust anything I backup to be successfully recoverable due to SCSI bus resets. We have made the following changes to our system.
Disable CDI
Disable the Dynamic Drive Sharing
Disable TURs by editing the Windows host¿s registry
Disable the Removable Storage Management (RSM) service
Ensure that Non-Rewind Devices are used
Single initiator / Single Target Zoning must be used.
ESM must be met for all Host Operating System patches, HBAs, SAN Switches, etc.
Conducted a CDL Volume Label Check
QLogic HBA enable target resets were disabled.

After making all of the changes above I personally relabeled all available volumes leaving me with only 7 days of backups. I marked all these volumes as read only, manual recycle. This would prevent me from writing to a potentially corrupt volume. I then over wrote my schedules to conduct a full backup, and backed up 404 servers. Today I ran scanner on 81 volumes. I have examined 65 of these 81 volumes and found 6 volumes corrupt still. I have worked this issue for about 10 days from when the SCSI bus reset was detected. I have clocked more than 200 hours in 9 days and this is still not resolved. It was detected by an event viewer report ID 9 and 11.

6 Operator

 • 

14.4K Posts

 • 

56.2K Points

May 18th, 2006 22:00

Do you have anything in SAN switch logs? If you don't have anymore overlapping zones that I would assume your problem comes from single point - try to discover on which host (storage node) was written. If you are able to do that and pinpoint single host then you focus on hardware on that box. Oh wait, if you have now no sharing and no overlapping zones then error should be visible in logs on that single box too... so can you isolate single system as source of this or do you see this on multiple storage nodes? If multiple, do those drives share same bus in the library? Do you see anything in library logs?

I'm not sure if your statement about SCSI resets being old only 9 days is accurate or not, but under assumption that is correct (timestamp when errors in event viewer appeared even Windows has crappy log system) are you aware of any changes to the system at that time. At the end it could even broken library. Your best approach to detect your problem is to force the error now by isolating piece by piece of HW until you find where it is. (or if you have enough money and you need urgent solution ask someone with SCSI/SAN analyzer/sniffer to give it a look - however that is extremely expensive).

40 Posts

May 19th, 2006 06:00

Do you have a case # for your corporately supported situation? That might be helpful to us

6 Operator

 • 

14.4K Posts

 • 

56.2K Points

May 22nd, 2006 06:00

Sorry, I'm a guy from EMC Software Group. We will
find out the errors and its helpfull, when you give
as your case# for more details. Thanks Fabio


I thought EMC logo should be next to the EMC people.

6 Operator

 • 

14.4K Posts

 • 

56.2K Points

May 22nd, 2006 07:00

Hi Fabio,

You should ask Documentum people running this forum. By default, each member of ESG should have sign next to the name which informs other users he belongs to ESG. At least that was design. You may ask this question to general supportability forum or perhaps forum forum/features at http://forums.emc.com/forums/forum.jspa?forumID=2

23 Posts

June 8th, 2006 14:00

Brian, "you kung fu no stwong" :-]

5 Posts

June 15th, 2006 17:00

Few things to try perhaps....
You may need to get emc to put an xgigx FC analyzer inline with the CDL to see when/why the resets or 'other' happen

If you have time try disabling specific OSs (to check if mixed OS+drivers play well together) or looking at the data on the tape with dd and compare to a good tape.(data shifted off 64kb boundaries?)

If you are using emulated IBM drives use the atdd driver

Try and copy a V tape to Real tape in the CDL console and then scan

There is a new version of CDL out 2.2 and networker 731 came out yesterday
Did you try the jumbo patch on 7.3?
No Events found!

Top