Announcement Banner

zainal1

updated

17 years ago

Z

zainal1

2 Intern

•

127 Posts

0

1083

August 3rd, 2009 17:00

why primary data-mover fail-over?

Hi

I am interested in finding out under what conditions does a primary DM
fail-over to the secondary DM. Besides hardware, what are the other issues
that can cause a DM to fail-over.

The reason I'm asking this is that my primary DM fail-over after a few weeks
of running. So far the secondary DM is running and has no issues.


Also what are the normal troubleshooting steps that a CE need to perform
on the Celerra to gather information.
  • gbarretoxx1

    2 Intern

    •

    366 Posts

    530

    0

    Posted August 6th, 2009 08:00

    Hi,

    you can check the /nas/var/dump directory for a text file that begins with "header".
    This is the dump header. If the panic created a memory dump, it would create also a text file with it's header.
    If you read this file, you would have an indication on what happened, but not all headers are "friendly", so the message might not be "readable" to you.
    An EMC resource should be able to identify what caused the issue, so I recommend you to open a support case, or ask your local representative if they received any case for this box.


    Gustavo Barreto.
  • dynamox

    11 Legend

    •

    20419 Posts

    •

    87439 Points

    530

    1

    Posted August 3rd, 2009 20:00

    Did the system call home ? I've had an instance where datamover was trying to read particular area on a file system and would panic. We ended up running full fsck on the file system to get things cleaned up.
  • zainal1

    2 Intern

    •

    127 Posts

    530

    0

    Posted August 3rd, 2009 20:00

    Yes, the system did call home.

    The NAS model is NS40. The current DART code is 5.6.37.6.
  • Rainer_EMC

    6 Operator

    •

    8645 Posts

    530

    0

    Posted August 4th, 2009 03:00

    Hi,

    the Configuring Standbys on EMC Celerra lists the concepts and the conditions that triger a data mover failover:

    ¿ Failure (operation below the configured threshold) of both internal network
    interfaces by the lack of a heartbeat (Data Mover time-out)
    ¿ Power failure within the Data Mover
    ¿ Software panic
    ¿ Exception on the Data Mover
    ¿ Data Mover hang
    ¿ Memory error on the Data Mover
    ¿ Removing a Data Mover from its slot (except on the NS20/NS40 series)

    and the conditions that do NOT cause a data mover failover:

    ¿ Manually restarting a Data Mover
    ¿ Removing a Data Mover from its slot (on the NS20/NS40/NX4 series only)

    The troubleshooting normally is looking at the log files to determine if the cause was hardware or software
  • Eclipsys1

    13 Posts

    530

    1

    Posted August 4th, 2009 15:00

    You may want to upgrade your code. We are upgraded to 5.6.44. There is supposedly a bug that causes the panic with versions earlier than .44. EMC could verify that.
  • dynamox

    11 Legend

    •

    20419 Posts

    •

    87439 Points

    530

    0

    Posted August 6th, 2009 19:00

    you can see how many systems are running that code:

    https://powerlink.emc.com/nsepn/webapps/btg548664833igtcuup4826/km/live1/en_US/Offering_Technical/Technical_Documentation/ActualCodes.pdf

    5.6.45.5 came out on July 2nd ..so it's relatively new. I am upgrading to it this Saturday as this is my last maintenance before December.
  • zainal1

    2 Intern

    •

    127 Posts

    530

    0

    Posted August 6th, 2009 19:00

    The "header" file gives this error mg.
    PANIC: I/O not progressing LastVol touched 86 Kind 0 (ptr=0x6952804)


    EMC CE recommended to upgrade the DART code to the latest which is 5.6.45.5.
    My current DART code is 5.6.37.6.

    How safe is 5.6.45.5? Is it recommended to use the latest code or we shld use at least 3 revs. below it.

    Thanks
  • Rainer_EMC

    6 Operator

    •

    8645 Posts

    530

    0

    Posted August 7th, 2009 03:00

    I would recommend to upgrade to 45