Announcement Banner
UNSOLVED

dcrawford21

updated

21 years ago

D

dcrawford21

17 Posts

0

4897

December 9th, 2005 10:00

Issues for a Networker client that has multiple network cards

Networker for Windows version 7.2.1 (both server & client)

Client:
MS Windows 2003 Server, Enterprise Edition, with SP1
SQL Server, verion 8.0 installed

This client has multiple NICs configured with HP teaming software for switch assisted load balancing with fault tolerance (SLB). I seem to remember seeing a solution, if not several, to this on the old support web page, but can no longer find it on this new support page.
The problem is, I can backup most of the data using the "ALL" save set, but one partition times out after about twelve hours of streaming data to tape (I watched it). If I break out the partition into multiple save sets and put it in a schedule by itself, the partition backs up OK, so this seems to be related to the lenth of time of the backup. The client is behind a stateful firewall, but the default ports for Networker have been opened in order to debug Networker (it fails completely with ports restricted).
Also, I cannot perform a manual backup from the client. I get an error saying that manual backups are not enabled on the Networker server, but I have verified that it is.
Any assistance would be much appreciated.

Thanks !
  • jtgowing

    28 Posts

    1543

    0

    Posted December 10th, 2005 03:00

    Hi,
    AFAIK teaming/trunking/ shouldn't affect Networker providing the client has only a single IP address/name.
    If you have implemented Multiple IP Interfaces, usually to provide a dedicated backup network, then you will need to specify entries in the Client\Preferences\Server Network Interface" and the Storage Node Affinity List."
    Specify the Name of the Networker Server that is on the desired Network, and the name of the Target Storage Node that is on the desired Network.

    Perhaps you could give more detail of your network setup, specifically does the Networker Server have one or more interfaces.

    The fact that most of your Backups are OK suggest the Network is OK. I think the issue is more to do with your problem Partition.
    What Size isit? Deos it have a very large filesystem with many directories and files?
    Do you have the problem when doing incrementals, but not with Fulls?

    A common cause of timeouts is when doing incremental backups of large filesystems with thousands/millions of files. The networker client can spend many minutes/hours searching massive filesystems for very few files which have changed, during this time the client does not send any data to the server and it can time out the client. You'll see an "Inactivity" time out in the Daemon Log.

    I am not aware that Manual backups can be enabled on disabled on the server, but remember that a manual backup does not autmatically get allocated to a pool, and so will default to the Default Pool, unless you explicitly specify a pool in your manual command.

    Lastly you will need to study the section in the Admin guide on setting up backup through firewalls. The process is not that difficult, but it is a bit tedious, and needs ome carefull information gathering and calculations, to establish the exact number and type ports to open. THere is also a tech bulletin on the subject on the Doc CD.

    Good Luck
  • ble1

    6 Operator

    •

    14354 Posts

    •

    56186 Points

    1543

    0

    Posted December 11th, 2005 08:00

    David,

    Let's start from the end.

    What have you verified? If NetWorker is complaining about manual backup chances are it is not enabled (this is not default). For example:
    C:\>save -s hcrvelin C:\Appl\tmp
    save: SYSTEM error: manual backups are not enabled on hcrvelin
    save: Cannot open save session with hcrvelin
    save completion time: 12-11-05 5:06p

    To get them back go into the server propeties and make sure that "Manual saves" is set to Enabled.

    Given the stream drops exactly after 12 hours would indicate TCP timeout setting within firewall or system OS itself. So speak to ppl managing that and make sure to check settings.

    In 7.1.x and after (and that included 7.2.x) Legato added NSR_KEEPALIVE_WAIT variable to address some FW issues they had with some FW settings. I believe before that you had to make sure that OS keepalive is below firewall's idle timeout for connection. Another manual workaround at that time was to run group in verbose mode (can be set from group properties). During the incremental backups or those that have to wait in the queue due to server parallelism you could see this issue. See official documentation for more information on the subject.

    I'm not aware that teaming software/setup should influence this anyhow, but checking setup properties it's always good thing to do.
  • dcrawford21

    17 Posts

    1548

    0

    Posted December 12th, 2005 09:00

    Thanks Hrvoje, guess this is a misprint in the 7.2.1 documentation, as there was no mention that it is expecting the value in seconds. Will give it a try.

    Thanks !
  • dcrawford21

    17 Posts

    1548

    0

    Posted December 12th, 2005 09:00

    Thanks for the reply John, please see my answers embedded in your text:

    Hi,
    AFAIK teaming/trunking/ shouldn't affect Networker providing the client has only a single IP address/name.

    Correct, the client is only using one IP address.

    If you have implemented Multiple IP Interfaces, usually to provide a dedicated backup network, then you will need to specify entries in the Client\Preferences\Server Network Interface" and the Storage Node Affinity List."
    Specify the Name of the Networker Server that is on the desired Network, and the name of the Target Storage Node that is on the desired Network.

    Perhaps you could give more detail of your network setup, specifically does the Networker Server have one or more interfaces.

    The Networker server has two NICs, each on it's own physical network. The appropriate entry was made in the client configurations for which NIC to use.

    The fact that most of your Backups are OK suggest the Network is OK. I think the issue is more to do with your problem Partition.
    What Size isit? Deos it have a very large filesystem with many directories and files?
    Do you have the problem when doing incrementals, but not with Fulls?

    The partition is large, about 100GB. It fails when I include it with the normal schedule, both incrementals and fulls. When I put it in it's own schedule and run it by itself, it passes as long as the partition is broken up into multiple save sets, but not as a single save set.

    A common cause of timeouts is when doing incremental backups of large filesystems with thousands/millions of files. The networker client can spend many minutes/hours searching massive filesystems for very few files which have changed, during this time the client does not send any data to the server and it can time out the client. You'll see an "Inactivity" time out in the Daemon Log.

    Yes, I am seeing this in the daemon log.

    I am not aware that Manual backups can be enabled on disabled on the server, but remember that a manual backup does not autmatically get allocated to a pool, and so will default to the Default Pool, unless you explicitly specify a pool in your manual command.

    Your right, didn't think about that, but would expect a different error message.

    Lastly you will need to study the section in the Admin guide on setting up backup through firewalls. The process is not that difficult, but it is a bit tedious, and needs ome carefull information gathering and calculations, to establish the exact number and type ports to open. THere is also a tech bulletin on the subject on the Doc CD.

    Been there, done that :) That's why we had the default ports used by Networker opened up temporarily for debugging.

    Thanks Again !
  • dcrawford21

    17 Posts

    1543

    0

    Posted December 12th, 2005 09:00

    I finally found the NSR_KEEPALIVE_WAIT variable documented in the 7.2.1 release notes, but there is not much on how to use it. Is it wanting seconds, minutes, or hours ? What would be an appropriate value for keeping a save set/group alive for twelve hours ?

    Thanks !
  • ble1

    6 Operator

    •

    14354 Posts

    •

    56186 Points

    1543

    0

    Posted December 12th, 2005 09:00

    You can find description in release supplement 7.1 as well - didn't check other instances. As it says, time is in seconds. You can try to set it up to 3500 and see what happens:

    "The period that nsrexecd will send keep alive messages to nsrexecd is adjustable by the NSR_KEEPALIVE_WAIT environment variable. Set this environment variable to the desired number of seconds between keep alive wait messages. If the environment variable is set to 0, a negative number or an invalid value or is not set, then no keep alive messages will be sent. The interval between keepalive messages may be slightly higher than the value set in NSR_KEEPALIVE_WAIT (it could be as much as 10 seconds off). This is necessary for code efficiency.

    The NSR_KEEPALIVE_WAIT variable sets the timeout limit that the nsrexecd daemon uses to keep messages active once connection to the NetWorker server has been established. If the NSR_KEEPALIVE_WAIT variable is not set or is set to an invalid value (0, a negative number, or a non-numeric string) then no keep alive message will be sent."
  • dcrawford21

    17 Posts

    1548

    0

    Posted December 14th, 2005 08:00

    OK, I put the save set back to ALL, but the e:\ partition failed again - "Aborted due to inactivity"...any other ideas ? Will investigate TCP settings on the firewall.
    Also, still not able to do manual backups from the client, even though I have verified that manual backups are enabled on the Networker server. It seems that this might be related to an authentication problem. The client is in an Active Directory domain, but the Networker server is in Workgroup mode. I tried adding the domain account to the admin group on the Networker server, but it does not seem to recognize it. I used the user=,domain= format.
    Guess I will have to open a case for this one. Thanks for the input, I'll post any progress.
  • ble1

    6 Operator

    •

    14354 Posts

    •

    56186 Points

    1548

    0

    Posted December 14th, 2005 09:00

    When you start the save from the client it has nothing to do in what domain it is - if you have path to that server backup will run. I will assume you are logged as admin user. Do following (on the client):
    - verify that nsrexecd is running as local system account
    - go to command line
    - run rpcinfo -p $your_backup_server_name_here
    - run rpcinfo -t $your_backup_server_name_here 390113 1
    - run tracert $your_backup_server_name_here

    Above should give you some meaningful results - if not you should trace the error.

    Further from command line:
    - run save -s $your_backup_server_name_here -b$your_pool_name_here -lfull -D9 E:

    Above should run save of E drive with level full to a pool you specify (make sure that level and pool are not excluding each other). If that fails debug output should give you more clue about where it is failing.

    Aborted due to inactivity here could also point to a disk problem - at least few times so far I came across where save could not read disk record (file system corruption or even disk had a bad sector) and running save in debug mode showed which file it was. When tried to manipulate manually with that file (eg. open it with explorer) it would fail too. If that is your case you should see in task manager save commands which are still hanging (unless you rebooted the client machine). It's worth to check it out.

    Also, while there is no reason to have inacitiry bigger than 30 minutes during full backup, during incremental it may happen. 30 minutes is default inside group setting so you may wish to increase it or for a quick test set it to 0 (which means no inactivity timeout setting) even that is not recommended in long term.


    Message was edited by: hcrvelin
  • jtgowing

    28 Posts

    1548

    0

    Posted December 15th, 2005 07:00

    Hi David,
    A Couple more comments:-

    I agree with Hrvoje, that the E:\ partition problem is looking like it may be disk related. Does Chkdsk or a more sophisticated diagnostic check out OK on that Partition?

    Just a little issue I've tripped over with Domain Member Clients and Non Domain Member servers (Win WorkGroup or Unix). When the client machine joins the domain it's machine name becomes fully qualified. If you have it specified by short name in Legato you MUST have the FQDN specified as an alias, and even then I've had grief. I make it a policy to always define windows domain clients by FQDN.
    You might want to try re-defining the client. (Remember to save the Client ID to transfer to the new one, to ensure access to your old savesets)
    At the risk of being boring, you must also check forward and reverse dns lookups, for short and FQDNames, from both server and client. A single anomoly can scupper the communications.

    Usually if it's a server client comms/authentication issue running a probe for the group for that client from the command line in verbose mode will often reveal a helpful error mesage.
    try savegrp -pvvv -l full -c "client" -G "Group"

    Good Luck
  • dcrawford21

    17 Posts

    391

    0

    Posted December 22nd, 2005 10:00

    Sorry, it took me a while to get back to this. The trace route works fine, but the rpcinfo -p from the either of the two clients gets the following:

    C:\>rpcinfo -p "Networker server hostname"
    program vers proto port
    100000 2 tcp 7938
    100000 2 udp 7938
    390113 1 tcp 7937
    390103 2 tcp 8826
    390109 2 tcp 8826
    390110 1 tcp 8826
    390120 1 tcp 8826
    390103 2 udp 9029
    390109 2 udp 9029
    390110 1 udp 9029
    390120 1 udp 9029
    390107 5 tcp 8335
    390107 6 tcp 8335
    390105 5 tcp 7976
    390105 6 tcp 7976
    390104 105 tcp 8191
    390104 205 tcp 9046
    390104 305 tcp 8778
    390104 405 tcp 9316
    rpcinfo: can't contact portmapper: Remote system error - Connection timed out

    The same command works fine when executed from the server side. If I execute this command from any of the other clients that are trouble free, I see this:

    program vers proto port
    100000 2 tcp 7938
    100000 2 udp 7938
    390113 1 tcp 7937
    390103 2 tcp 9857
    390109 2 tcp 9857
    390110 1 tcp 9857
    390120 1 tcp 9857
    390103 2 udp 8591
    390109 2 udp 8591
    390110 1 udp 8591
    390120 1 udp 8591
    390107 5 tcp 8146
    390107 6 tcp 8146
    390105 5 tcp 9824
    390105 6 tcp 9824
    390104 105 tcp 9129
    390104 205 tcp 9500
    390104 305 tcp 9907
    390104 405 tcp 8804
    program vers proto port
    100000 2 tcp 7938
    100000 2 udp 7938
    390113 1 tcp 7937
    390103 2 tcp 9857
    390109 2 tcp 9857
    390110 1 tcp 9857
    390120 1 tcp 9857
    390103 2 udp 8591
    390109 2 udp 8591
    390110 1 udp 8591
    390120 1 udp 8591
    390107 5 tcp 8146
    390107 6 tcp 8146
    390105 5 tcp 9824
    390105 6 tcp 9824
    390104 105 tcp 9129
    390104 205 tcp 9500
    390104 305 tcp 9907
    390104 405 tcp 8804

    So it's looking like a portmapper issue on the client side. Any ideas on how I should proceed. By the way, I should mention that these two servers use Citrix, but don't know any further details since I know very little about it.

    Thanks In Advance !