I just upgraded from RM4 to RM5.0.2. My old application sets and jobs were not as clean as I would have liked so I deleted everything down to the storage and recreated the application sets, storage pools, jobs, and schedules (no big deal, I only have a dozen or so).
SnapView snaps are working great.
Some of my SanCopy incremental jobs work fine, but two of them (one is NTFS source, the other is SQL2000) consisently error like this:
2007 06 02 12:17:15 mail03b 000033 ERROR: Function dcSANCopyMarkSessions failed.
2007 06 02 12:17:15 mail03b 000618 ERROR: Replication Manager was unable to unmark the Incremental SAN Copy session clariion_RMINC-APM00012345678_0030-061116130145 on array APM00012345678.
Note that the initial SanCopy incremental session went without error.
All the hosts (those that are working and those that are having issues) are Windows 2003 Enterprise Edition SP1.
Yes we did upgrade to Flare 24 back in Feb 07 (as well as all the host based software to .24; admsnap, navicli, naviagent,etc) based on recommendation from Support to resolve a queued I/O problem we were having on the CX.
I took a look at emc161580 and we are running Solutions Enabler 6.4.0.5 on all the hosts, so I don't think that is related even though the log entries look similar. I may try Solutions Enabler 6.3.2.20 on one of the hosts just too see if that makes a difference.
If you think of anything else, please let me know. Otherwise it's off to Support!
Funnily enough I just had a call with this exact error for consistently failing Sancopy jobs after upgrading to RM 5.0.2 (and upgrading navi agent/cli, etc). All hosts except this host was working for Incr Sancopy.
I fixed (at least this particular instance of the error) it by addding in the required entries in the Navisphere agent's agent.config file. The host had a blank agent.config. Typically C:\Program Files\emc\navisphere agent\agent.config I added the required lines, eg
user (username)@spA user (username)@spB user system@spA user system@spB
(The user will need to be the same as all the other users in the RM configuration to ensure the clariion authorization works.)
Restart Naviagent service and then Replication Manager Client service. Try your Incr sancopy job again.
Got exact same issue. Running FLARE 24 and upgraded to RM5 Sp2. Tried agent config file, and solution en 6.3.2.20 - no luck. The odd thing is that jobs are failing RANDOMLY. Meaning the same job may fail and then work again.
We had another customer who had this issue and they opened a call with us and engineering have identified the problem in the RM Code, which is due to the new Flare 24 single management interface design.
As long as your issue can be confirmed to be the same, there is a hotfix available for you. If you open a Service Request and ask for Hotfix 31704 for RM 5.0.2, it will resolve your problem.
If you search using the Windows Search option; for "files containing the following words" in the RM Client debug logs directory on the production or local mount host (C:\program Files\emc\rm\logs\client\)
and find the following; main. CLAR Err: Error returned from Agent main. CLAR Err: This command was sent to the SP that does not own the lun (0x71008043)
Then you can confirm it is the same issue and the hotfix will resolve this.
This is also happening to my setup after a new RM 5.0.2 install. I am waiting for the HotFix # 31704 to fix the way the naviseccli interprets the owner information in the listsessions output. If you have the jobs retry about 3 times in 60 second intervals, they will eventually grab the proper SP and the jobs will complete. This shouldn't be a permanent fix but will get you through till you receive the HotFix.
JamesBEMC
257 Posts
325
1
Posted June 5th, 2007 05:00
Is there any pattern to the failures? ie only Windows 2000 hosts affected?
Did you also upgrade the Flare code on the Clariion?
You may want to read the EMC KnowledgeBase solution emc161580 on Powerlink.
If you are still in trouble, please open an EMC SR and we'll investigate for you.
Cheers
James