UNSOLVED

john_s_main

updated

14 years ago

J

john_s_main

132 Posts

0

136778

August 14th, 2012 13:00

vFoglight alerts to run

Here are the alerts which we think are important, the rest we turn off:

  • CPU % Utilization – various measures – Fatal level only 
    • Cluster CPU Utilized > 60%
    • ESX CPU Utilized > 70%
  • VM % ready > 5%
  • VM and ESX Memory Swapping – any is bad, not the same as OS level swapping, this is swapping at the ESX host level, causes massive performance issues
  • ESX or VM Memory Balloon Deflation Failure (Balloon memory not being released, indicates an OS level issue)
  • Cluster vRAM allocated > 100%
  • Cluster vRAM Utilized > 70%
  • Low Datastore Disk Space - Cluster datastore with less than 50GB free is a real issue when you have to do a restore
  • Low VM Volume disk space – Fatal level only – need to decide on thresholds (C:\ Drive with less than (vRAM Amount) space free, for example)
  • VMWare Tools not running/reporting - only send alerts once per day
  • ESX Hosts disconnected from VirtualCenter
  • ESX Host multipathing outage
  • Agent Data Collection Alerts - working on some self-healing capabilities for these
  • FMS Operational Alerts
  • Rule to clear any alert after 3 days - assuming you don't have people actively clearing them (shocking, I know)

I'm looking for others' suggestions for important things to monitor.