Announcement Banner
UNSOLVED

Sturm3

updated

17 years ago

S

Sturm3

27 Posts

0

2068

January 13th, 2010 02:00

Tape Drive Stucked

Hey,

I am using Networker 7.3 with eml 245 tape library (6 drives). Firstly one of my tape drives got stucked at "verify label.." operation. After some search I have found that this is because of not using persisting binding of wwn. I have changed it but it needs a reboot. The problem is that i cant reboot the library because monthly backups have started and one of them is about 11 TB and pretty slow. I dont trust to the restart option because once it didnt work and I have to backup whole again. I have also tried to empty all tape drives with hardware and "restart library" option via software gui or console but failed. One more drive is now stucked also at the moment. I know that this is a software issue. Can anyone help? Thanks.

  • masonb

    445 Posts

    906

    1

    Posted January 13th, 2010 03:00

    There are a few things to consider here.  You say you have configured Persistent binding but need a library reboot to clear the "stuck tape".  You do not say if you have restarted the NetWorker software on the storage node concerned or NetWorker server.  Once verifying label is reported its only really cleared by restarting the nsrmmd.  If the drive has issues after this you can sometimes clear this via other commands, but when the library comes back up it checks the status of the drives and becomes ready in NetWorker.  So you could need ot restart the software to clear the nsrmmd or the library has reported the drive as OK but the tape is still in the drive somehow.If the tape volume is physically in the drive we can attempt to get it back to the slot, but this may not necessarily make the drive available in NetWorker.So you can do the following: -1.  Does nsrjb output for the jukebox report the volume in the drive?  If yes then NW knows the tape is there and if drive is reporting verify label you will have to stop the software on the applicable storage node which maybe you do not want to do as there are backups running to other drives which are OK.  If its the NetWorker server you will not be able to do anything until your backups have finished.2.  Can you disable and enable the drive in NetWorker? - if not you will have to wait until backups have completed before doing anything.3.  Check the status of the drives directly with the jukebox (not NetWorker) with sjirdtag b.t.l | more command.  The sji and mt binaries/executables are supplied with the NetWorker code but are not maintained by EMC. Its standard SCSI code for interfacing with robotics/drives/scsi devices.b.t.l is the scsi bus, target, lun number of the robotic arm (you can find this by inquire command or usually reported in the jukebox resource unless you have changed the default name.  Example output below, here drive 1 in the jukebox has tape with barcode U00114 in it and the other drive is not loaded.  Slots 1-3, 5-9 have tapes in them, 4, 10 and 11 do not have tapes in.  You should be able to correlate with NetWorker nsrjb output where tapes are loaded and not: -C:\Documents and Settings\bmason>sjirdtag 7.1.1Tag Data for 7.1.1, Element Type DATA TRANSPORT:Elem[001]: tag_val=1 pres_val=1 med_pres=1 med_side=0VolumeTag= Elem[002]: tag_val=0 pres_val=1 med_pres=0 med_side=0Tag Data for 7.1.1, Element Type STORAGE:Elem[001]: tag_val=1 pres_val=1 med_pres=1 med_side=0VolumeTag= Elem[002]: tag_val=1 pres_val=1 med_pres=1 med_side=0VolumeTag= Elem[003]: tag_val=1 pres_val=1 med_pres=1 med_side=0VolumeTag= Elem[004]: tag_val=0 pres_val=1 med_pres=0 med_side=0Elem[005]: tag_val=1 pres_val=1 med_pres=1 med_side=0VolumeTag= Elem[006]: tag_val=1 pres_val=1 med_pres=1 med_side=0VolumeTag= Elem[007]: tag_val=1 pres_val=1 med_pres=1 med_side=0VolumeTag= Elem[008]: tag_val=1 pres_val=1 med_pres=1 med_side=0VolumeTag= Elem[009]: tag_val=1 pres_val=1 med_pres=1 med_side=0VolumeTag= Elem[010]: tag_val=0 pres_val=1 med_pres=0 med_side=0Elem[011]: tag_val=0 pres_val=1 med_pres=0 med_side=04.  Look for the barcode reported in the first stanza under DATA TRANSPORT:  if it not reported there then the tape should not be in the drive.  If it is you can use the following to eject and move the tape from the drive.mt -f offline - on the storage nodeon the server which controls the robotic arm for the jukebox: -sjimm b.t.l drive xxx slot yyy - where xxx is the Elem number of the drive reported in sjirdtag output and yyy is the slot reported in NW or an empty slot within the jukebox. e.g.sjimm 7.1.1 drive 001 slot 004 would move the tape from drive 1 to slot 4 providing you have run the mt command on the correct OS device otherwise the command will fail as the tape drive will not have released the tape ready for a move. Please be aware you need to ensure you are working on the correct drive as these commands are outside of Networker knowledge and if you specify the wrong one and its being used or a backup you will compromise anything writing to the drive.I would also disable and enable the drive within NetWorker after doing this, this should be allowed by the software and set the software for future operations..Regards,Bill
  • Sturm3

    27 Posts

    906

    0

    Posted January 13th, 2010 04:00

    Hey William, thanks for the answer.

    1- I couldnt restart neither the server,library nor the nsr service. So persisted binding isn't applied yet, reboot of the server is required I think.

    2- By the way I forget to mention that I am using "hp command view tl" web interface. It is the web interface of the library itself.That interface can move drives, list the status of the drivers... Is it the same as your code?

    3- I cant disable the stucked devices. All I can do is to make them at service. At least I dont have to wait for timeouts in loading/unloading.

    4- Inquire command stalls. It shows part of the devices and stops responding.

    nsrjb shows the devices which has problem as empty

    web interface shows the devices which has problem as empty

    networker gui shows the devices which has problem as full

    I have tried to unload all of the tape drives by web interface. It succeded. Networker didnt update the status of the drives normally as it writes "hardware move will make backups fail and you need to reinventory software". I tried to inventory via nw but failed. Message was something like there is no media in the drive. Finally I tried the reset option via nw. It seems like it succeded, updated the status of working drives but halt when it come to the stucked drive.

    So i need to wait until backups finish? Thanks...


  • masonb

    445 Posts

    906

    0

    Posted January 13th, 2010 06:00

    OK - so it sounds like you have two issues here: -

    1. Most important is the persistent binding - you need to ge this sorted and configured.  There is a very good document with a lot of information on what to do and useful info like CDI settings etc.  Search on Powerlink for "Configuring Tape Devices for EMC NetWorker" - each time you use check the version is current as it is maintained by EMC.

    2. NMC GUI refresh has nto worked correctly.

    If nsrjb and your vendor GUI show the drive as empty but NMC does not, that should not be an issue and you can wait for your backups to complete.  Once complete I would ensure your drives/jukeboxes and OS devices are correct according to the document advice making any server reboots as necessary.  You may have to reconfigure the NetWorker jukebox if the device paths changed afterwards but that is simple and quick compared to the other work to get everything right.  From then on it should be plain sailing!

    The Vendor GUI probably runs the commands I specified without you knowing (or something very similar) - you can use these instead of mt and sjirdtag/sjimm.

    Regards,

    Bill