First, I would go back to the original memory configuration that was booting correctly before the upgrade. Then confirm that the server starts without memory errors.
On our side there is also a fatal PCIe error on the Tesla T4 in slot 1, I would also try the same test without the GPU installed, just to isolate the issue properly.
Once the server is stable, I would perform a full power drain / flea drain following Dell’s procedure:
Then I would test the new DIMMs by replacing existing known-good DIMMs in the same slots, without changing the overall memory population at first. This should confirm whether the new memory modules are working correctly.
We actually had a similar case at a customer site yesterday. We followed this approach — removing the GPU, performing a proper power drain, then reinstalling and testing the new DIMMs progressively — and the server came back up without memory errors.
If the new DIMMs work fine as replacements, then I would retry the complete upgrade while strictly following the Dell / VxRail memory population rules.
The goal is to avoid combining too many variables at once: new DIMMs, new slots, GPU/PCIe error, firmware state, and previous POST/MRC memory errors.
Hi, In fact, there may be some problems when using non-Dell-certified parts. That doesn't mean those parts won't work, but they may not work with full functionality or cause other problems. If I would have done the necessary checks according to the Memory guidelines and I am getting an error. If Dell-certified DIMM was used here, if the BIOS and iDRAC are up to date, I would check with Memory cross test to see if the problem is in the slot or in the memory itself. I would check to see if different DIMMs on the same slot give the same error. If necessary, I would do the minimum and check the new processor with a DIMM that I'm sure is working. I would be the last to suspect the motherboard and processor, in case there is a problem with the memory channel.
Also, I want to share a useful article "DDR4 memory: managing Correctable error threshold events" https://dell.to/3KY5tDt
DELL-Erman O
Social Media and Communities Professional
Dell Technologies | Enterprise Support Services
#IWork4Dell
Did I answer your query? Please click on ‘Mark as Accepted Answer’. ‘Thumbs up’ the posts you like!
I have the same problem in the same Server R650xs 2X4316, adding extra 6X32 GB 3200 RAM (Micron and supposed to be DELL Original) to the standard 2X32 GB 3200, total is 8X32; I have the same errors.
Could you let us know if the BIOS and iDRAC/LCC have been updated. What are the slots that you have installed the memories in? How many processor configurations is in the server? May I know if the memories are RDIMM? How many ranks are they? Have you tried to install just the new memories, does the error prompts?
DELL-Joey C
Social Media and Communities Professional
Dell Technologies | Enterprise Support Services
#IWork4Dell
Did I answer your query? Please click on ‘Mark as Accepted Answer’ if I did.
UnD3R2C0R3
1 Rookie
•
1 Message
•
5 Points
0
0
Posted June 11th, 2026 18:51
Hello,
I would troubleshoot this step by step.
First, I would go back to the original memory configuration that was booting correctly before the upgrade. Then confirm that the server starts without memory errors.
On our side there is also a fatal PCIe error on the Tesla T4 in slot 1, I would also try the same test without the GPU installed, just to isolate the issue properly.
Once the server is stable, I would perform a full power drain / flea drain following Dell’s procedure:
https://www.dell.com/support/kbdoc/en-us/000175625/how-do-i-reset-and-drain-power-of-my-dell-poweredge-server
Then I would test the new DIMMs by replacing existing known-good DIMMs in the same slots, without changing the overall memory population at first. This should confirm whether the new memory modules are working correctly.
We actually had a similar case at a customer site yesterday. We followed this approach — removing the GPU, performing a proper power drain, then reinstalling and testing the new DIMMs progressively — and the server came back up without memory errors.
If the new DIMMs work fine as replacements, then I would retry the complete upgrade while strictly following the Dell / VxRail memory population rules.
The goal is to avoid combining too many variables at once: new DIMMs, new slots, GPU/PCIe error, firmware state, and previous POST/MRC memory errors.
Best regards,