We have a C8000 enclosure populated with eight C8220 sleds, each fitted with dual Intel E5-2680 2.7 GHz CPUs and 128 Gb of installed memory. One of these sleds however has always been slow - it takes hours to install any Linux OS onto it from a USB DVD drive when all the others installed quickly and in operation, it is slow to respond to commands or run compute jobs.
Since the original BIOS 1.1.11 was reporting the CPUs were running at 3500MHz instead of 2700, a BIOS update to the latest 2.3 version was done and has fixed this mis-reporting. The server is still slow and looking in the output from dmesg, I found these entries whch differ from all the other sleds:
Hierarchical RCU implementation. [ 0.000000] RCU dyntick-idle grace-period acceleration is enabled. [ 0.000000] RCU restricting CPUs from NR_CPUS=256 to nr_cpu_ids=120. [ 0.000000] Offload RCU callbacks from all CPUs [ 0.000000] Offload RCU callbacks from CPUs: 0-119.
It is possible that there is a CPU issue. Before swapping or replacing CPUs I would check to see if throttling is occurring. Check the chassis to see if power usage is anywhere near the cap. If a PSU fails or power usage meets/exceeds the capabilities of the PSUs then throttling can occur.
I would also test the node that is having issues in another slot before swapping/replacing hardware.
Sorry for the late reply, had to schedule downtime on another sled so I could try your suggestion of moving the defective node to another bay in the chassis. This has had no effect on the problem I'm afraid.
I have also checked the power and throttling settings in the BIOS and they are identical to the other nodes. I don't think throttling can be an issue really because this node is sluggish even at idle. We are running Ubuntu Server 14.04 on these nodes but CentOS 6.6 has also been tried with the same results.
I think I will book a support call as this node is under warranty.
Daniel My
12 Elder
•
6205 Posts
284
0
Posted July 1st, 2015 13:00
Hello Andy
It is possible that there is a CPU issue. Before swapping or replacing CPUs I would check to see if throttling is occurring. Check the chassis to see if power usage is anywhere near the cap. If a PSU fails or power usage meets/exceeds the capabilities of the PSUs then throttling can occur.
I would also test the node that is having issues in another slot before swapping/replacing hardware.
Thanks