Unsolved

This post is more than 5 years old

7 Posts

46088

March 18th, 2014 06:00

Odd planar temp warnings - OMSA

One of our Dell servers (PowerEdge R210 II, running OpenSUSE) has recently started to report abnormal alerts for the system board planar temperature.

OMSA logs show series of warning/critical messages alerting of planar temp getting above the upper limits and then (within few seconds/minutes) getting beyond the lower limits:

Critical       - Tue Mar 18 09:29:22 2014 - The system board planar temperature is greater than the upper critical threshold.
Non-critical - Tue Mar 18 09:29:22 2014 - The system board planar temperature is greater than the upper warning threshold.
Normal       - Tue Mar 18 09:29:27 2014 - The system board planar temperature is within range.
Non-critical - Tue Mar 18 09:29:27 2014 - The system board planar temperature is greater than the upper warning threshold.
Non-critical - Tue Mar 18 09:30:22 2014 - The system board planar temperature is less than the lower warning threshold.
Normal       - Tue Mar 18 09:30:28 2014 - The system board planar temperature is within range. -
Non-critical - Tue Mar 18 09:30:38 2014 - The system board planar temperature is less than the lower warning threshold.

We have been checking ambiental and chassis temp by hand, as well as other indicators (fans, temps, voltages), and everything seems to be ok within ranges.

Since lower and upper thresholds are set to 8ºC and 42ºC (~46ºF to ~107ºF) and it is obvious that the server is not suffering such extreme rises and falls of temperature, I guess either the temp probe is failing or maybe there is some problem with the firmware.

Has anyone run into this problem? Any advice or suggestions will be greatly appreciated.

Community Manager

 • 

9.7K Posts

 • 

43.2K Points

March 18th, 2014 08:00

Hi,

What version is the bios and BMC/iDRAC firmware at? Since it is showing in the logs it is probably not OMSA causing it. Has the system been rebooted?

7 Posts

March 18th, 2014 09:00

Hi Josh,

The system was rebooted just a week ago. We noticed these alerts some days before this scheduled reboot, and they keep on appearing ever since.

The firmware info:

BIOS Information
         Vendor: Dell Inc.
         Version: 1.0.3
         Release Date: 04/01/2011
         BIOS Revision: 1.0

BMC version: 1.70.00 (Build 6)

Cheers

Community Manager

 • 

9.7K Posts

 • 

43.2K Points

March 18th, 2014 10:00

The BIOS is very far out of date and the BMC is slightly out of date, updating these may set the sensors back to a normal state so that they stop incorrectly reporting the temperature. Since you are running opensuse and there is not an in OS upgrade method, the easiest method is to use our liveDVD and then run the redhat updates from there.

 

BIOS: http://www.dell.com/support/drivers/us/en/04/DriverDetails/Product/poweredge-r210-2?driverId=H2P3G&osCode=ES11&fileId=3348639009&languageCode=EN&categoryId=BI

 

BMC: http://www.dell.com/support/drivers/us/en/04/DriverDetails/Product/poweredge-r210-2?driverId=F4D3G&osCode=ES11&fileId=3163533232&languageCode=EN&categoryId=ES

 

LiveDVD http://linux.dell.com/files/openmanage-contributions/omsa-65-live/OMSA65-CentOS6-x86_64-LiveDVD.iso

7 Posts

March 19th, 2014 05:00

Thanks for your quick reply!

we will proceed with these updates as soon as possible (we have to first move some services to other servers).

Once we complete these updates I will post how it works out.

Thanks again!

7 Posts

March 26th, 2014 04:00

We successfully updated the BIOS and the BMC two days ago, but the system keeps reporting the same odd temperature alerts, sometimes above the upper thresholds, sometimes below the lower thresholds.

Any further ideas or suggestions?

Community Manager

 • 

9.7K Posts

 • 

43.2K Points

March 26th, 2014 08:00

You could try resetting the NVRAM on the BIOS with the motherboard jumper. Page 115 ftp://ftp.dell.com/Manuals/all-products/esuprt_ser_stor_net/esuprt_poweredge/poweredge-r210-2_Owner%27s%20Manual_en-us.pdf

 

7 Posts

April 7th, 2014 11:00

It's been some days since last post. Just a recap of what we have done so far.

We reset the NVRAM, rebooted the system and kept it "under supervision" for some days. The first couple of days no temp alerts were reported, but then the "warning" and "critical" messages appeared again. Just a few messages the very next days, but after a couple more days the number of alerts grew really fast. Eventually, the log filled up and we had to clear it off.

We then repeated these steps again, reseting the NVRAM and monitoring the server. And again, the first days after this reset the system has not reported any new temp alert, but now these messages have reappeared in the log.

I'd greatly appreciate any other suggestions.

Community Manager

 • 

9.7K Posts

 • 

43.2K Points

April 7th, 2014 11:00

Is the system under warranty? At this point it sounds like the sensors are having some issue.

7 Posts

April 8th, 2014 02:00

I am afraid that warranty has already expired.

I assume that these sensors are embedded in the motherboard. If so, what are the alternatives now? Replacing the motherboard?

Community Manager

 • 

9.7K Posts

 • 

43.2K Points

April 8th, 2014 08:00

Yes, replacing the motherboard would be the way to replace the sensors.

7 Posts

April 8th, 2014 09:00

Not the best news...

Just one more question. Are server motherboards available for ordering through Dell site? I have not found any suitable motherboard on the spare parts section.

Thank you very much for your help these days.

Community Manager

 • 

9.7K Posts

 • 

43.2K Points

April 8th, 2014 09:00

They may not be on the website, you may need to call our parts department, 800-357-3355

No Events found!

Top