UNSOLVED

Disk Jockey

updated

17 years ago

DJ

Disk Jockey

1 Rookie

•

92 Posts

0

454

August 19th, 2009 19:00

Two Celerras panicked at the same time...

NS960 integrated and NS40G both panic at very nearly the same time last night.
Both at 5.6.45-5
Arrays are fine.
Both on different ethernet switches, one on fabric, one not, as it is integrated.
Backups had finished well prior to crash.
Usual CIFS file activity going on.

Common infrastructure:
CAVA 4.2.2 and 5.2.3.29, 4 cavas shared between Celerras.
Netbackup 6.0 MP5 NDMP, not sharing drives between these two units.
No complete dumps, support is looking into the logs.

Severe thunderstorms in NYC last night, but all of this gear is on the same circuits as the arrays, backed by a 120KVA Liebert which did not kick in at all.

Very puzzling..

From our friends at Symantec:
Symantec AV Engine Rollback ¿ 8/19/09

Symantec has received a limited number of reports of system instability
after applying the new AV Engine (released in multiple daily definitions and
rapid release definitions of August 18th rev 16 or beyond). At this time
reports are confined to Vista systems running an older version of SEP 11.0
(MR 2).

This update was not released in certified definitions to systems running SAV
or SCS. As a precaution, Symantec is releasing a definition update and the
engine version will revert to 20081.3.2.13 beginning with August 19th, rev
21 definitions. Additional notification will be sent when a new AV Engine
will be released.
Message was edited by:
spaceman
  • nandas

    6 Operator

    •

    1473 Posts

    198

    0

    Posted August 20th, 2009 07:00

    There are many different reason for which a data mover may panic. It is very difficult to comment not looking at the panic headers.

    And when you say the NS960 or NS40G panicked - did you mean all the data movers in each box panicked? How many data movers are there in each box?

    Anyway - I assume, you have the EMC Dial home setup for both the boxes - the panic would have created dial home cases and EMC Support must be working on the same. You may request them to find out the root cause for the panic which will provide the details - provided the panic dumps are available without which it may not be possible for EMC Support team to do the analysis. They cna tell you more on this.

    Thanks,
    Sandip
  • Disk Jockey

    1 Rookie

    •

    92 Posts

    198

    0

    Posted August 20th, 2009 08:00

    Yes, dialouts worked well and I was called immediately by support. They are looking at the logs.

    Bottom line-two completely independent NAS heads crash at the same time-what are the chances of that?

    DM2 in an NS40 and DM2 in an NS960 each panicked at the same time (very closely) 9:20pm, August 18th.

    Bottom line-two completely independent NAS heads crash at the same time-what are the chances of that?

    2DM in each box, so they both failed over properly and service continued on the standbys.

    These boxes have been rock solid. I find it hard to blame EMC.

    Symantec's notice yesterday of scan engine instability caused by an update on the 18th leads me to look in their direction. Leave it to Symantec to destabilize just about any system....

    Message was edited by:
    spaceman