Unsolved

This post is more than 5 years old

4 Posts

6257

March 29th, 2007 11:00

I/O Device Error: Bad Tape, Bad Drive, or Bad Networker?

Running Legato Networker 7.2.2 J on Windows 2003. Recently upgraded from 7.1.3.

Seeing issues in our daemon.log indicating that I/O errors are aborting backups or are finishing tapes long before they've reached capacity. This is resulting in several of our larger save sets to go un-backed up without multiple retries. This issue is spanning multiple servers of various purposes, but seems to be tied mostly to larger save sets (200g+).

This issue started, more or less, after the upgrade. Could be a coincidence. Doesn't seem to be consistent with tape numbers (some are much older, some are fairly new). Doesn't seem to be consistent with drives (this example spans two different ones).

Below is a daemon.log sample. There are others related to this that are occurring fairly regularly. I figured I'd check here to get an opinion as to source before I escalated to EMC Support.

Thank you for any help our advice you can provide!

------------------------------------------------------------------

03/29/07 09:43:13 nsrd: media warning: \\.\Tape0 writing: The request could not be performed because of an I/O device error., at file 35 record 25505
03/29/07 09:43:13 nsrd: media notice: LTO Ultrium-2 tape XXXX on \\.\Tape0 is full
03/29/07 09:43:13 nsrd: media notice: LTO Ultrium-2 tape XXXX used 70 GB of 190 GB capacity
03/29/07 09:43:13 nsrd: media warning: verification of volume "XXXX", volid 3388920884 failed, read open error: drive status is Drive reports no error - but state is unknown
03/29/07 09:43:13 nsrd: media notice: verification of volume "XXXX", volid 3388920884 failed, volume is being marked as full.
03/29/07 09:43:13 nsrmmd #30: Diagnostic: remember_as_high: no available sop for ssid 3658139812!
03/29/07 09:43:13 nsrd: write completion notice: Writing to volume XXXX complete
03/29/07 09:43:13 nsrd: media notice: Save set (3658139812) client1.domain.com:F:\ volume XXXX on \\.\Tape0 is being terminated because: Media verification failed
03/29/07 09:43:13 nsrd: client1.domain.com:F:\ done saving to pool 'POOL' (XXXX) 276 GB
03/29/07 09:43:40 savegrp: command 'save -s 10.231.4.13 -g POOLGrp -LL -m client1.domain.com -l full -q -W 78 -N F:\ F:\ ' for client client1.domain.com exited with return code 255.
03/29/07 09:43:40 savegrp: client1.domain.com:F:\ will retry 1 more time(s)
03/29/07 09:43:42 nsrd: media info: suggest mounting XXYY on tapeserver.domain.com for writing to pool 'POOL'
03/29/07 09:43:42 nsrd: media waiting event: Waiting for 1 writable volumes to backup pool 'POOL' tape(s) on tapeserver.domain.com
03/29/07 09:43:43 nsrd: \\.\Tape3 Eject operation in progress
03/29/07 09:44:13 nsrd: media info: loading volume XXYY into \\.\Tape3
03/29/07 09:44:25 nsrd: \\.\Tape3 Verify label operation in progress
03/29/07 09:44:40 nsrd: \\.\Tape3 Mount operation in progress
03/29/07 09:44:54 nsrd: media event cleared: Waiting for 1 writable volumes to backup pool 'POOL' tape(s) on tapeserver.domain.com
03/29/07 09:44:54 nsrd: client1.domain.com:F:\ saving to pool 'POOL' (XXYY)
03/29/07 09:53:57 nsrd: media warning: \\.\Tape3 writing: The request could not be performed because of an I/O device error., at file 2 record 24247
03/29/07 09:53:57 nsrd: media notice: LTO Ultrium-2 tape XXYY on \\.\Tape3 is full
03/29/07 09:53:57 nsrd: media notice: LTO Ultrium-2 tape XXYY used 1551 MB of 190 GB capacity
03/29/07 09:53:57 nsrd: media info: verification of volume "XXYY", volid 3372143768 succeeded.
03/29/07 09:53:57 nsrd: write completion notice: Writing to volume XXYY complete
03/29/07 09:53:57 nsrd: media info: suggest mounting XXXY on tapeserver.domain.com for writing to pool 'POOL'
03/29/07 09:53:57 nsrd: media waiting event: Waiting for 1 writable volumes to backup pool 'POOL' tape(s) on tapeserver.domain.com
03/29/07 09:53:58 nsrd: \\.\Tape3 Eject operation in progress
03/29/07 09:54:26 nsrd: media info: loading volume XXXY into \\.\Tape3

4 Posts

March 29th, 2007 12:00

No, this spans across all the tape drives and has probably spanned across three dozen tapes. The tapes span a whole range of usage and purchase date. Some of which were bought as recently as last year with 3 or fewer uses.

Driver could be likely, but, again, this issue didn't appear to be occuring in my previous Legato installation using the same driver.

6 Operator

 • 

14.4K Posts

 • 

56.2K Points

March 29th, 2007 12:00

AFAIK CDI works fine with LTO2 so that should not be an issue. If you want you can try to downgrade to 7.1.4 to see if you still get this. However, the truth is that EOT gets reported by driver to NW. You could try to play with this if you get this on daily basis (eg. for each tape) and test it with NT backup.

4 Posts

March 29th, 2007 12:00

I went backed and checked my Daemon.log files from my old version (again, 7.1.3) and I can find no instance at all of these type of errors. I see I/O errors from day one of the 7.2.2 J install.

Is it possible that 7.1.3 just wasn't capturing this data before?

6 Operator

 • 

14.4K Posts

 • 

56.2K Points

March 29th, 2007 12:00

I doubt it's bad NetWorker.

If it is bad drive then you would get these always on the same drive. Do you?

Bad tape - well, if you get more of those that could be bad tapes, but I would always suspect more SAN or tape driver.

35 Posts

March 30th, 2007 02:00

The I/O errors are physical errors, this may be a hardware problem (HBA; drives or volumes) but not about the NetWorker software, try to mount and unmount the damaged volumes and if they fails is cause tne volume is physical damage, if not it can be a problem with the drive or the HBA, switch, see the system event logs of the backup server for errors and search in the daemon,log if the problem occurs always with the same drive or with the same volumes

69 Posts

March 30th, 2007 06:00

This is interesting. we are a fairly new installation (7.2.1) and have seen the same issues. We are running Windows 2003, Qlogic 2340 HBA's in Dell servers (6650 and 2850). the driver is certified by EMC for Legato and SAN.

One recent fix we applied is to ensure all dedicated and storage node systems have persistent binding set on the HBA. I deleted and redefine all device definitions after this was set/reset. (jbedit)

After this was done i had removed the media (file and media index entries) and relable the tapes. I used the reconfigured device during the relabel to ensure connectivitey to the device was functional. A pain... but once completed this problem has disappeared.

4 Posts

April 1st, 2007 16:00

Your response makes me think that it IS Networker.

I can confirm that it's not the driver at this stage. I've used two different driver versions since I posted this with the same result.

The errors are always occuring on the SAME save sets as well. Specifically 5 or so larger save sets.

I can confirm that persistant binding is set correctly on the HBA but I may blow it away and try again.

I'm also going to escalate this to Legato as well because several of our server backups are outright failing and this is causing me one large headache.

6 Operator

 • 

14.4K Posts

 • 

56.2K Points

April 2nd, 2007 01:00

Have you test it with NT backup?

1 Message

February 1st, 2012 00:00

Sorry to bring up an old thread. However we are getting the same errors on large savegroups following an upgrade to NW 7.6.2.

We have 3 different server devices using the same tape library and only one of these seems to be producing the errors.

Tape library has 18 drives and picks tapes from a pool of 250 for these particular jobs. I would discount it being media error or device issues as the problem seems to occur no matter which tapes are used and across various drives. It also repeatedly happens on the same save groups, which leads me to believe its likely to be the large jobs spanning across multiple tapes that dont work.

It does seem to be occuring when the tape fills up. It seems that the EOT is not recognised by NW and it just keeps attempting to send data to it.

The strange thing is, if a new savegroup is created with just the failed parts of the jobs it works successfully when run the morning after it fails.

I will be deleting and recreating the problem jobs today to see if that addresses the issue, However i was hoping someone who may have originally posted on this topic found a solution to their issues?

No Events found!

Top