Unsolved
This post is more than 5 years old
4 Posts
1
6257
March 29th, 2007 11:00
I/O Device Error: Bad Tape, Bad Drive, or Bad Networker?
Running Legato Networker 7.2.2 J on Windows 2003. Recently upgraded from 7.1.3.
Seeing issues in our daemon.log indicating that I/O errors are aborting backups or are finishing tapes long before they've reached capacity. This is resulting in several of our larger save sets to go un-backed up without multiple retries. This issue is spanning multiple servers of various purposes, but seems to be tied mostly to larger save sets (200g+).
This issue started, more or less, after the upgrade. Could be a coincidence. Doesn't seem to be consistent with tape numbers (some are much older, some are fairly new). Doesn't seem to be consistent with drives (this example spans two different ones).
Below is a daemon.log sample. There are others related to this that are occurring fairly regularly. I figured I'd check here to get an opinion as to source before I escalated to EMC Support.
Thank you for any help our advice you can provide!
------------------------------------------------------------------
03/29/07 09:43:13 nsrd: media warning: \\.\Tape0 writing: The request could not be performed because of an I/O device error., at file 35 record 25505
03/29/07 09:43:13 nsrd: media notice: LTO Ultrium-2 tape XXXX on \\.\Tape0 is full
03/29/07 09:43:13 nsrd: media notice: LTO Ultrium-2 tape XXXX used 70 GB of 190 GB capacity
03/29/07 09:43:13 nsrd: media warning: verification of volume "XXXX", volid 3388920884 failed, read open error: drive status is Drive reports no error - but state is unknown
03/29/07 09:43:13 nsrd: media notice: verification of volume "XXXX", volid 3388920884 failed, volume is being marked as full.
03/29/07 09:43:13 nsrmmd #30: Diagnostic: remember_as_high: no available sop for ssid 3658139812!
03/29/07 09:43:13 nsrd: write completion notice: Writing to volume XXXX complete
03/29/07 09:43:13 nsrd: media notice: Save set (3658139812) client1.domain.com:F:\ volume XXXX on \\.\Tape0 is being terminated because: Media verification failed
03/29/07 09:43:13 nsrd: client1.domain.com:F:\ done saving to pool 'POOL' (XXXX) 276 GB
03/29/07 09:43:40 savegrp: command 'save -s 10.231.4.13 -g POOLGrp -LL -m client1.domain.com -l full -q -W 78 -N F:\ F:\ ' for client client1.domain.com exited with return code 255.
03/29/07 09:43:40 savegrp: client1.domain.com:F:\ will retry 1 more time(s)
03/29/07 09:43:42 nsrd: media info: suggest mounting XXYY on tapeserver.domain.com for writing to pool 'POOL'
03/29/07 09:43:42 nsrd: media waiting event: Waiting for 1 writable volumes to backup pool 'POOL' tape(s) on tapeserver.domain.com
03/29/07 09:43:43 nsrd: \\.\Tape3 Eject operation in progress
03/29/07 09:44:13 nsrd: media info: loading volume XXYY into \\.\Tape3
03/29/07 09:44:25 nsrd: \\.\Tape3 Verify label operation in progress
03/29/07 09:44:40 nsrd: \\.\Tape3 Mount operation in progress
03/29/07 09:44:54 nsrd: media event cleared: Waiting for 1 writable volumes to backup pool 'POOL' tape(s) on tapeserver.domain.com
03/29/07 09:44:54 nsrd: client1.domain.com:F:\ saving to pool 'POOL' (XXYY)
03/29/07 09:53:57 nsrd: media warning: \\.\Tape3 writing: The request could not be performed because of an I/O device error., at file 2 record 24247
03/29/07 09:53:57 nsrd: media notice: LTO Ultrium-2 tape XXYY on \\.\Tape3 is full
03/29/07 09:53:57 nsrd: media notice: LTO Ultrium-2 tape XXYY used 1551 MB of 190 GB capacity
03/29/07 09:53:57 nsrd: media info: verification of volume "XXYY", volid 3372143768 succeeded.
03/29/07 09:53:57 nsrd: write completion notice: Writing to volume XXYY complete
03/29/07 09:53:57 nsrd: media info: suggest mounting XXXY on tapeserver.domain.com for writing to pool 'POOL'
03/29/07 09:53:57 nsrd: media waiting event: Waiting for 1 writable volumes to backup pool 'POOL' tape(s) on tapeserver.domain.com
03/29/07 09:53:58 nsrd: \\.\Tape3 Eject operation in progress
03/29/07 09:54:26 nsrd: media info: loading volume XXXY into \\.\Tape3
Seeing issues in our daemon.log indicating that I/O errors are aborting backups or are finishing tapes long before they've reached capacity. This is resulting in several of our larger save sets to go un-backed up without multiple retries. This issue is spanning multiple servers of various purposes, but seems to be tied mostly to larger save sets (200g+).
This issue started, more or less, after the upgrade. Could be a coincidence. Doesn't seem to be consistent with tape numbers (some are much older, some are fairly new). Doesn't seem to be consistent with drives (this example spans two different ones).
Below is a daemon.log sample. There are others related to this that are occurring fairly regularly. I figured I'd check here to get an opinion as to source before I escalated to EMC Support.
Thank you for any help our advice you can provide!
------------------------------------------------------------------
03/29/07 09:43:13 nsrd: media warning: \\.\Tape0 writing: The request could not be performed because of an I/O device error., at file 35 record 25505
03/29/07 09:43:13 nsrd: media notice: LTO Ultrium-2 tape XXXX on \\.\Tape0 is full
03/29/07 09:43:13 nsrd: media notice: LTO Ultrium-2 tape XXXX used 70 GB of 190 GB capacity
03/29/07 09:43:13 nsrd: media warning: verification of volume "XXXX", volid 3388920884 failed, read open error: drive status is Drive reports no error - but state is unknown
03/29/07 09:43:13 nsrd: media notice: verification of volume "XXXX", volid 3388920884 failed, volume is being marked as full.
03/29/07 09:43:13 nsrmmd #30: Diagnostic: remember_as_high: no available sop for ssid 3658139812!
03/29/07 09:43:13 nsrd: write completion notice: Writing to volume XXXX complete
03/29/07 09:43:13 nsrd: media notice: Save set (3658139812) client1.domain.com:F:\ volume XXXX on \\.\Tape0 is being terminated because: Media verification failed
03/29/07 09:43:13 nsrd: client1.domain.com:F:\ done saving to pool 'POOL' (XXXX) 276 GB
03/29/07 09:43:40 savegrp: command 'save -s 10.231.4.13 -g POOLGrp -LL -m client1.domain.com -l full -q -W 78 -N F:\ F:\ ' for client client1.domain.com exited with return code 255.
03/29/07 09:43:40 savegrp: client1.domain.com:F:\ will retry 1 more time(s)
03/29/07 09:43:42 nsrd: media info: suggest mounting XXYY on tapeserver.domain.com for writing to pool 'POOL'
03/29/07 09:43:42 nsrd: media waiting event: Waiting for 1 writable volumes to backup pool 'POOL' tape(s) on tapeserver.domain.com
03/29/07 09:43:43 nsrd: \\.\Tape3 Eject operation in progress
03/29/07 09:44:13 nsrd: media info: loading volume XXYY into \\.\Tape3
03/29/07 09:44:25 nsrd: \\.\Tape3 Verify label operation in progress
03/29/07 09:44:40 nsrd: \\.\Tape3 Mount operation in progress
03/29/07 09:44:54 nsrd: media event cleared: Waiting for 1 writable volumes to backup pool 'POOL' tape(s) on tapeserver.domain.com
03/29/07 09:44:54 nsrd: client1.domain.com:F:\ saving to pool 'POOL' (XXYY)
03/29/07 09:53:57 nsrd: media warning: \\.\Tape3 writing: The request could not be performed because of an I/O device error., at file 2 record 24247
03/29/07 09:53:57 nsrd: media notice: LTO Ultrium-2 tape XXYY on \\.\Tape3 is full
03/29/07 09:53:57 nsrd: media notice: LTO Ultrium-2 tape XXYY used 1551 MB of 190 GB capacity
03/29/07 09:53:57 nsrd: media info: verification of volume "XXYY", volid 3372143768 succeeded.
03/29/07 09:53:57 nsrd: write completion notice: Writing to volume XXYY complete
03/29/07 09:53:57 nsrd: media info: suggest mounting XXXY on tapeserver.domain.com for writing to pool 'POOL'
03/29/07 09:53:57 nsrd: media waiting event: Waiting for 1 writable volumes to backup pool 'POOL' tape(s) on tapeserver.domain.com
03/29/07 09:53:58 nsrd: \\.\Tape3 Eject operation in progress
03/29/07 09:54:26 nsrd: media info: loading volume XXXY into \\.\Tape3
No Events found!


AL2099
4 Posts
0
March 29th, 2007 12:00
Driver could be likely, but, again, this issue didn't appear to be occuring in my previous Legato installation using the same driver.
ble1
6 Operator
•
14.4K Posts
•
56.2K Points
0
March 29th, 2007 12:00
AL2099
4 Posts
0
March 29th, 2007 12:00
Is it possible that 7.1.3 just wasn't capturing this data before?
ble1
6 Operator
•
14.4K Posts
•
56.2K Points
0
March 29th, 2007 12:00
If it is bad drive then you would get these always on the same drive. Do you?
Bad tape - well, if you get more of those that could be bad tapes, but I would always suspect more SAN or tape driver.
EEHO
35 Posts
0
March 30th, 2007 02:00
dmorcio
69 Posts
0
March 30th, 2007 06:00
One recent fix we applied is to ensure all dedicated and storage node systems have persistent binding set on the HBA. I deleted and redefine all device definitions after this was set/reset. (jbedit)
After this was done i had removed the media (file and media index entries) and relable the tapes. I used the reconfigured device during the relabel to ensure connectivitey to the device was functional. A pain... but once completed this problem has disappeared.
AL2099
4 Posts
0
April 1st, 2007 16:00
I can confirm that it's not the driver at this stage. I've used two different driver versions since I posted this with the same result.
The errors are always occuring on the SAME save sets as well. Specifically 5 or so larger save sets.
I can confirm that persistant binding is set correctly on the HBA but I may blow it away and try again.
I'm also going to escalate this to Legato as well because several of our server backups are outright failing and this is causing me one large headache.
ble1
6 Operator
•
14.4K Posts
•
56.2K Points
0
April 2nd, 2007 01:00
hardke01
1 Message
0
February 1st, 2012 00:00
Sorry to bring up an old thread. However we are getting the same errors on large savegroups following an upgrade to NW 7.6.2.
We have 3 different server devices using the same tape library and only one of these seems to be producing the errors.
Tape library has 18 drives and picks tapes from a pool of 250 for these particular jobs. I would discount it being media error or device issues as the problem seems to occur no matter which tapes are used and across various drives. It also repeatedly happens on the same save groups, which leads me to believe its likely to be the large jobs spanning across multiple tapes that dont work.
It does seem to be occuring when the tape fills up. It seems that the EOT is not recognised by NW and it just keeps attempting to send data to it.
The strange thing is, if a new savegroup is created with just the failed parts of the jobs it works successfully when run the morning after it fails.
I will be deleting and recreating the problem jobs today to see if that addresses the issue, However i was hoping someone who may have originally posted on this topic found a solution to their issues?