UNSOLVED

Jodi-2015

updated

20 years ago

J2

Jodi-2015

17 Posts

0

1674

November 27th, 2006 06:00

daemons hanging

Good Morning,

Our current setup is a Networker server running 7.3.2 on a Solaris 8 platform with storage node running 7.3.2 on a Solaris 9 platform. Almost all clients are running the same version. We have a few cluster clients where the physical node has been upgraded, but the virtual node still reports the old version.

Several times our server has just hung. At first there were GSS authentication errors reported. We disabled GSS and the same thing happened this weekend. Three backups were running and just stopped. There were no errors reported, below is the section of the daemon.log:

11/26/06 02:38:07 nsrd: stacker3:index:drnotes1 done saving to pool 'IndxSt2Weekly' (A05038) 502 MB
11/26/06 02:38:27 nsrd: stacker3:index:bhnsm02 saving to pool 'Weekly' (A06435L2)
11/26/06 02:43:03 savegrp: job (53779) host: f68vrec savepoint: /u01 had WARNING indication(s) at completion.
pools supported: IndxSt2Weekly;
11/26/06 03:50:09 savegrp: Aborting inactive job (53513) lnotes3:/plaza2
11/26/06 03:50:10 savegrp: command 'save -s stacker3 -g fullweekly -LL -f - -m lnotes3 -l full -q -W 78 -N /plaza2 /plaza2 ' for client lnotes3 exited with return code 1.
11/26/06 03:50:10 savegrp: job (53513) host: lnotes3 savepoint: /plaza2 had WARNING indication(s) at completion.
11/26/06 03:50:10 savegrp: lnotes3:/plaza2 failed.
* lnotes3:/plaza2 (interrupted), exiting
* lnotes3:/plaza2 aborted due to inactivity
11/26/06 03:50:10 savegrp: lnotes3:/plaza2 will retry 1 more time(s)
11/26/06 04:03:05 nsrlcpd #1: Jukebox `lto1' is exiting. The jukebox is no longer managed by nsrlcpd.

Then there was nothing in the daemon log and nsrwatch got hung as well. I could not even do an nsr_shutdown. I had to kill all processes one at a time. Then when I restarted the daemons, there was an error saying the volumes have reached the low water mark. A second restart of the daemons cleared that.

Any ideas?