I have changed the BIOS setting (after I remembered it the first time booting failed with four modules) and I can boot the server with two of these installed (in A1 and A2) - the BIOS says that the RAM is running at 800MHz and I understand that to be a limitation of the CPU. But if I install all four modules, or even just three (in A1, A2, A3) the server appears not to boot - it sits at "Configuring memory, please wait." I haven't left it for hours to get past that, could it be that that's what I need to do?
I'm wondering if this might also be a limitation of the processor. But the processor spec page seems to indicate that it supports three memory channels and much more than the 64GB I have.
Some clarifications:
I took out the modules which were in A1 and A2 and put in there the ones which were in A3 and A4. The machine still boots. So it does not seem that the issue is a faulty module.
All firmware is at the latest version, to the best of my knowledge - with the exception of the BMC (BIOS is at 1.14).
It turns out that my previous messages were on the right track. There was damage to the CPU socket, fortunately not of the replace-the-motherboard variety.
I ordered a pair of CPUs (as I thought I could at least use all four DIMMs by putting two on each CPU). When I installed the first CPU I noticed something not quite right in the socket but ignored it at that point. Here's what I saw (I have circled the problem):
But with the second CPU installed I was getting E1410 errors and the machine wouldn't boot so I had to take everything apart again and investigate.
When I looked more closely I realised that this pin was out of place (I'm not sure if it was bent or just caught on the neighbouring pin) but I managed to get it back from the red position to the green position.
That didn't solve my dual-CPU issue but I found that with one CPU, it would now see all 64GB, which it had not done before. So the original memory issue was apparently due to that pin above.
Ultimately (for anyone who is interested) I did resolve the E1410 error - I think what did it was cleaning the second CPU socket by dabbing with a tissue moistened with isopropyl alcohol as Intel or Dell documentation mentions somewhere. As of now the machine boots with two L5640s and 64GB RAM (16GB in each of A1, A2, B1, B2), and I hope it will remain stable.
The first thing I would do is to verify if the server is up to date on BIOS, iDrac, etc. After that would you confirm the specific part number for the dimms, as it may be down to a compatibility issue as well?
Let me know.
DELL-Chris H
Social Media and Communities Professional
Dell Technologies | Enterprise Support Services
#IWork4Dell
Did I answer your query? Please click on ‘Mark as Accepted Answer’. ‘Thumbs up’ the posts you like!
The Lifecycle Controller Platform Update Current Version screen reports the following (not in this order):
BIOS 1.14.0
iDRAC6 2.92
Dell LifeCycle Controller 1.7.5.4 X02
Dell 32 Bit Diagnostics 5162A0
OS Drivers 0
Broadcom NetXtreme II Gigabit Ethernet [2 of these, embedded] 7.12.19
SAS9211-8i [This is a PERC H310 flashed to IT mode] (Slot 4-1) 00.00.00.00
The memory modules are from Micron and marked as follows: MT36KSF2G72PZ-1G6E1HI 16GB 2RX4 PC3L-12800R-11-13-E2 From this Micron document I understand the last part to reflect module type R (most likely Registered?) CAS Latency 11, JEDEC SPD revision 13, Reference-raw-card-and-revision E2.
It seems that I have three modules from one batch and one from another: 3 x Build Lot ID DPAEEHY016 Date code 1438 1 x Build Lot ID DPAE8H9009 Date code 1430 But I have difficulty imagining that that alone would cause a problem (and the machine does boot with at least one combination of one of the first three and the one from the second lot).
I will try to do a careful visual check of slots 3 and 4. Maybe they have some sort of damage.
But so long as "supported" doesn't mean "the only things our hardware will agree to work with", surely not being on the list doesn't mean they definitely won't work and there's no point in looking for a solution...?
What exactly is happening during that "configuring memory" stage? Because at this point all I can tell is that it's there that adding 1-2 more modules makes things worse. Maybe if I can get past that all will be well? Some years ago you responded to a similar question and there apparently the server moved on eventually.
But thank you for the list of supported parts. Maybe I will check if I can return these modules and I'll order 20D6F which seems to be available. They're more expensive but at least theoretically are meant to work. So long as we don't know what the problem is now, it's hard to be sure that it won't affect supported modules too.
Unless the following images are faked, or the modules are actually made with some change specifically for Dell which makes them work differently (I find both hard to believe) then it seems that my modules are the same as the modules which are sold as 20D6F...
They are all MT36KSF2G72PZ-1G6E1 - the two characters which follow apparently indicate differences in components; mine are different from these but they vary between the Dell modules too.
So it seems to me that I should be able to consider these modules effectively supported, even if Dell won't actually commit to anything.
The machine boots with modules 1 and 2 in slots A1 and A2, and with modules 3 and 4 in slots A1 and A2. I haven't tried each module alone in A1, and that seems unnecessary.
But I haven't tried all combinations of two in A1 and A2; maybe I will find that some combination (like 2 and 3) doesn't let it boot; perhaps both of those were present when I had three modules installed.
It's rather annoying when these obscure problems crop up, but I always tell people who are frustrated by scenarios like these that no-one promised that every piece of technology which "should" work will work, and certainly not every time.
Any further ideas will be welcome. Is there any chance that this is related to the BMC? That may well still be at an early version.
of course we always suggest to update all firmware revision of the components, so if BMC is not updated, I invite you to update it and see if it resolve.
Thanks
Marco
DELL- Marco B
Social Media and Communities Professional
Dell Technologies | Enterprise Support Services
#IWork4Dell
Did I answer your query? Please click on ‘Mark as Accepted Answer’. ‘Thumbs up’ the posts you like!
ylavi
1 Rookie
•
15 Posts
658
0
Posted December 19th, 2022 15:00
It turns out that my previous messages were on the right track. There was damage to the CPU socket, fortunately not of the replace-the-motherboard variety.
I ordered a pair of CPUs (as I thought I could at least use all four DIMMs by putting two on each CPU). When I installed the first CPU I noticed something not quite right in the socket but ignored it at that point. Here's what I saw (I have circled the problem):
But with the second CPU installed I was getting E1410 errors and the machine wouldn't boot so I had to take everything apart again and investigate.
When I looked more closely I realised that this pin was out of place (I'm not sure if it was bent or just caught on the neighbouring pin) but I managed to get it back from the red position to the green position.
That didn't solve my dual-CPU issue but I found that with one CPU, it would now see all 64GB, which it had not done before. So the original memory issue was apparently due to that pin above.
Ultimately (for anyone who is interested) I did resolve the E1410 error - I think what did it was cleaning the second CPU socket by dabbing with a tissue moistened with isopropyl alcohol as Intel or Dell documentation mentions somewhere. As of now the machine boots with two L5640s and 64GB RAM (16GB in each of A1, A2, B1, B2), and I hope it will remain stable.