Unsolved
This post is more than 5 years old
40 Posts
0
2559
May 18th, 2006 21:00
SCSI bus resets driving me crazy
Hello all knowledgeable NetWorker Admins. I have been experiencing SCSI bus resets for the past 9 days. My environment is very unstable and I cannot trust anything I backup to be successfully recoverable due to SCSI bus resets. We have made the following changes to our system.
Disable CDI
Disable the Dynamic Drive Sharing
Disable TURs by editing the Windows host¿s registry
Disable the Removable Storage Management (RSM) service
Ensure that Non-Rewind Devices are used
Single initiator / Single Target Zoning must be used.
ESM must be met for all Host Operating System patches, HBAs, SAN Switches, etc.
Conducted a CDL Volume Label Check
QLogic HBA enable target resets were disabled.
After making all of the changes above I personally relabeled all available volumes leaving me with only 7 days of backups. I marked all these volumes as read only, manual recycle. This would prevent me from writing to a potentially corrupt volume. I then over wrote my schedules to conduct a full backup, and backed up 404 servers. Today I ran scanner on 81 volumes. I have examined 65 of these 81 volumes and found 6 volumes corrupt still. I have worked this issue for about 10 days from when the SCSI bus reset was detected. I have clocked more than 200 hours in 9 days and this is still not resolved. It was detected by an event viewer report ID 9 and 11.
Disable CDI
Disable the Dynamic Drive Sharing
Disable TURs by editing the Windows host¿s registry
Disable the Removable Storage Management (RSM) service
Ensure that Non-Rewind Devices are used
Single initiator / Single Target Zoning must be used.
ESM must be met for all Host Operating System patches, HBAs, SAN Switches, etc.
Conducted a CDL Volume Label Check
QLogic HBA enable target resets were disabled.
After making all of the changes above I personally relabeled all available volumes leaving me with only 7 days of backups. I marked all these volumes as read only, manual recycle. This would prevent me from writing to a potentially corrupt volume. I then over wrote my schedules to conduct a full backup, and backed up 404 servers. Today I ran scanner on 81 volumes. I have examined 65 of these 81 volumes and found 6 volumes corrupt still. I have worked this issue for about 10 days from when the SCSI bus reset was detected. I have clocked more than 200 hours in 9 days and this is still not resolved. It was detected by an event viewer report ID 9 and 11.
No Events found!


ble1
6 Operator
•
14.4K Posts
•
56.2K Points
0
May 18th, 2006 22:00
I'm not sure if your statement about SCSI resets being old only 9 days is accurate or not, but under assumption that is correct (timestamp when errors in event viewer appeared even Windows has crappy log system) are you aware of any changes to the system at that time. At the end it could even broken library. Your best approach to detect your problem is to force the error now by isolating piece by piece of HW until you find where it is. (or if you have enough money and you need urgent solution ask someone with SCSI/SAN analyzer/sniffer to give it a look - however that is extremely expensive).
brianluck1
40 Posts
0
May 19th, 2006 06:00
ble1
6 Operator
•
14.4K Posts
•
56.2K Points
0
May 22nd, 2006 06:00
find out the errors and its helpfull, when you give
as your case# for more details. Thanks Fabio
I thought EMC logo should be next to the EMC people.
ble1
6 Operator
•
14.4K Posts
•
56.2K Points
0
May 22nd, 2006 07:00
You should ask Documentum people running this forum. By default, each member of ESG should have sign next to the name which informs other users he belongs to ESG. At least that was design. You may ask this question to general supportability forum or perhaps forum forum/features at http://forums.emc.com/forums/forum.jspa?forumID=2
cluu1
23 Posts
0
June 8th, 2006 14:00
rutherfc
5 Posts
0
June 15th, 2006 17:00
You may need to get emc to put an xgigx FC analyzer inline with the CDL to see when/why the resets or 'other' happen
If you have time try disabling specific OSs (to check if mixed OS+drivers play well together) or looking at the data on the tape with dd and compare to a good tape.(data shifted off 64kb boundaries?)
If you are using emulated IBM drives use the atdd driver
Try and copy a V tape to Real tape in the CDL console and then scan
There is a new version of CDL out 2.2 and networker 731 came out yesterday
Did you try the jumbo patch on 7.3?