I am interested in finding out under what conditions does a primary DM
fail-over to the secondary DM. Besides hardware, what are the other issues
that can cause a DM to fail-over.
The reason I'm asking this is that my primary DM fail-over after a few weeks
of running. So far the secondary DM is running and has no issues.
Also what are the normal troubleshooting steps that a CE need to perform
on the Celerra to gather information.
you can check the /nas/var/dump directory for a text file that begins with "header". This is the dump header. If the panic created a memory dump, it would create also a text file with it's header. If you read this file, you would have an indication on what happened, but not all headers are "friendly", so the message might not be "readable" to you. An EMC resource should be able to identify what caused the issue, so I recommend you to open a support case, or ask your local representative if they received any case for this box.
Did the system call home ? I've had an instance where datamover was trying to read particular area on a file system and would panic. We ended up running full fsck on the file system to get things cleaned up.
the Configuring Standbys on EMC Celerra lists the concepts and the conditions that triger a data mover failover:
¿ Failure (operation below the configured threshold) of both internal network interfaces by the lack of a heartbeat (Data Mover time-out) ¿ Power failure within the Data Mover ¿ Software panic ¿ Exception on the Data Mover ¿ Data Mover hang ¿ Memory error on the Data Mover ¿ Removing a Data Mover from its slot (except on the NS20/NS40 series)
and the conditions that do NOT cause a data mover failover:
¿ Manually restarting a Data Mover ¿ Removing a Data Mover from its slot (on the NS20/NS40/NX4 series only)
The troubleshooting normally is looking at the log files to determine if the cause was hardware or software
You may want to upgrade your code. We are upgraded to 5.6.44. There is supposedly a bug that causes the panic with versions earlier than .44. EMC could verify that.
gbarretoxx1
2 Intern
•
366 Posts
530
0
Posted August 6th, 2009 08:00
you can check the /nas/var/dump directory for a text file that begins with "header".
This is the dump header. If the panic created a memory dump, it would create also a text file with it's header.
If you read this file, you would have an indication on what happened, but not all headers are "friendly", so the message might not be "readable" to you.
An EMC resource should be able to identify what caused the issue, so I recommend you to open a support case, or ask your local representative if they received any case for this box.
Gustavo Barreto.