I have a NW server with staging configure.
I notice on Sunday staging process stuck and client session remain idle, the disk was full, and the automatic staging not work.
Also I try manually to stage manually from the interface but without any success.
I check the processes with ps and I notice that are a lot of nsrindex process in the memory
It is not clear what happened here except that staging process was affected. Given that client sessions were stuck I would assume something was going rather than staging problem. Try good old one - stop everzthing and restart it. If it happens again check the logs and if you find nothing open this with support.
I did see recently (7.1.3b421) where fulll file system would not trigger staging. The issue why it happened in the first place was that file system check window was big enough to allow device to be filled.. thus leading to the problem where device would get fillled. Once that it was filled it did not stage back anything. So, I'm not sure if this was one off only or there is indeed issue with device when it reaches 100%.. What you could do is to list ssids for the give volumes and stage them manually. That's what I did and afterwards everything was fine.
To be more clear The disk pool use by adv_filedevice was 100% full above the watermark of 60%, and the automatically stage doesn¿t work. Is normal that the client remain in idle because didn¿t have anymore space left to direct the backup. Normally the staging process should be work. This is very strange.
To be more clear the NW server is install on Redhat. I have another questions, do you encounter situation when mminfo not respond. I encounter this situation last week before adv_file stuck
I have another questions, do you encounter situation when mminfo not respond. I encounter this situation last week before adv_file stuck
Nsrstage to stage the data needs to obtain data about ssids which are matching the criteria set in NSR staging resource to be staged. This is done by querying media database (mminfo or some other process doing it at internal level). I would assume if there is a problem in media db responding then nsrstage is going to be only one of the processes affected.
When mminfo dose't responde I found the following in my log, see below :
03/24/06 01:51:42 nsrd: Calling mm_deactivate for mmd 1 thats using device /staging/backup with volume Disk.001 on host null 03/24/06 01:51:42 nsrd: write completion notice: Writing to volume Disk.001 complete 03/24/06 01:51:44 nsrd: re-setting overall number of authorized save sessions, changing value from (1) to (0) 03/24/06 01:54:46 nsrmmd #1: Device /staging/backup/_AF_readonly, with total 82569883 blocks, each of size 4096, now has 42997403 blocks free 03/24/06 02:04:46 nsrmmd #2: Device /staging/backup/_AF_readonly, with total 82569883 blocks, each of size 4096, now has 42997403 blocks free 03/24/06 02:04:50 nsrd: deactivating mmd #1 03/24/06 02:04:50 nsrd: Calling mm_deactivate for mmd 1 thats using device /staging/backup with volume Disk.001 on host null 03/24/06 02:04:50 nsrd: write completion notice: Writing to volume Disk.001 complete 03/24/06 02:04:52 nsrd: re-setting overall number of authorized save sessions, changing value from (1) to (0) 03/24/06 02:14:46 nsrmmd #2: Device /staging/backup/_AF_readonly, with total 82569883 blocks, each of size 4096, now has 42997403 blocks free 03/24/06 02:18:16 nsrd: deactivating mmd #1 03/24/06 02:18:16 nsrd: Calling mm_deactivate for mmd 1 thats using device /staging/backup with volume Disk.001 on host null 03/24/06 02:18:16 nsrd: write completion notice: Writing to volume Disk.001 complete 03/24/06 02:18:18 nsrd: re-setting overall number of authorized save sessions, changing value from (1) to (0) 03/24/06 02:24:47 nsrmmd #1: Device /staging/backup/_AF_readonly, with total 82569883 blocks, each of size 4096, now has 42997403 blocks free 03/24/06 02:31:57 nsrd: deactivating mmd #1 03/24/06 02:31:57 nsrd: Calling mm_deactivate for mmd 1 thats using device /staging/backup with volume Disk.001 on host null 03/24/06 02:31:57 nsrd: write completion notice: Writing to volume Disk.001 complete
There is nothing to worry from the output you sent. All looks fine. Not sure how you figure out mminfo would be dead during that time, but it could be that your server is busy and that response is kind of slow during that time.
Yes, the server is very busy an load is betwen 1,3 and 2,5.Also are a lot of savegrp in the memory, and the mminfo display nothing.After I stop the nsr service, the load remain high. I have notice that the load decrease only when I kill the savegrp proceses from the memory. Next time I will wait to see if mminfo display something. Also I setup the paralelism like in Networker Admin Guide.
ble1
6 Operator
•
14354 Posts
•
56186 Points
443
0
Posted March 27th, 2006 00:00