UNSOLVED

ble1

updated

12 years ago

B

ble1

6 Operator

•

14354 Posts

•

56186 Points

0

2457

September 18th, 2014 13:00

5.4.2.1

Does anyone have experience with this DDOS release?  We recently updated 4 small DD boxes (DD160s and DD620s) and we observe following:

- if you export /ddvar and try to access it you are getting NFS stale mount messages if you do ls or cd into any of folders

- replication (CCR with NW) does fail on random basis with "Failed to perform DDCL filecopy: Starting a file copy failed ([5005] no room left)." which doesn't make sense as there are no disk capacity issues

- getting SBU is almost impossible as when you wish to download it doesn't give latest one or open main page again (and via CLI you get NFS message).

I opened a ticket with support partner today, but no meaningful response yet and I was wondering if anyone else has seen this. My worry now is that something might have gone wrong during upgrade (I found that 5.4.3.1 has fix for some timeout issue during upgrade, but these are all small systems so I doubt we have this issue).

  • Leo Li

    4 Apprentice

    •

    9030 Posts

    839

    0

    Posted September 22nd, 2014 20:00

    Hello,

    I would recommend you to please contact EMC support team to work on this issue. Thanks.

  • ble1

    6 Operator

    •

    14354 Posts

    •

    56186 Points

    839

    0

    Posted September 23rd, 2014 04:00

    They are already on it... once we know for sure what is going on, I will post resolution.

  • cmartinjr

    40 Posts

    839

    0

    Posted October 6th, 2014 05:00

    Odd, I'm running into the same issue and found this thread by a google search.  I'm running 5.4.0.8 and am running into the issue.  I would think it's networker, but I have two different datazones and both are just randomly pushing out this error on my cloning.  I'm going to go the data domain route and open up a ticket with support on it.

  • ble1

    6 Operator

    •

    14354 Posts

    •

    56186 Points

    839

    0

    Posted October 6th, 2014 06:00

    This is actually NW issue IMHO.  I will post a blog tonight on this topic.  As workaround create /nsr/debug/nsrcloneconfig with following inside (on each storage node including backup server):

    max_threads_per_client=0

    max_client_threads=0

    You will need to restart NW for this to take place.

  • cmartinjr

    40 Posts

    839

    0

    Posted October 6th, 2014 06:00

    Normally I would think it's a networker issue, but I have 2 different datazones (one running 8.1.1.8, the other running 8.0.2.1).  The one running 8.0.2 has been running that version for a while now and I just recently started running into the "no room left" error.

    I clone by group (from a clone script) and have it set to where no more than 20 groups (10 from one storage node, 10 from another) are cloning at the same time in my 8.1 environment, which is the larger of the two.

    I did search for the "max_threads_per_client" and found an article where it talks about parallel cloning at:

    NetWorker 8.0 SP4 is released! | NetWorker Lounge

    Is this what you're referring to?  Did adding those variables resolve the issue you were running into?

  • ble1

    6 Operator

    •

    14354 Posts

    •

    56186 Points

    839

    0

    Posted October 6th, 2014 10:00

    Yes and yes (that's why I told you to use it)

  • cmartinjr

    40 Posts

    839

    0

    Posted October 6th, 2014 11:00

    Cool.. 'preciate the info..

    I'm also exceeding the max number of read/write streams on my DD860.  DD support showed me in the ddfs.info log where it shows I'm exceeding this number which, according to them, will cause it to drop the extra streams.  It's not a constant that I'm over the max streams so all of my cloning is not failing, but it does happen at the start of my cloning process.

  • ble1

    6 Operator

    •

    14354 Posts

    •

    56186 Points

    839

    0

    Posted October 6th, 2014 15:00

  • cmartinjr

    40 Posts

    839

    0

    Posted October 7th, 2014 14:00

    That is an awesome writeup on the issue!  Thank you, I understand exactly what's going on now..