I am facing an issue where the machine reboots in case the pcie link goes down. In my case the PCIE device is an FPGA which goes down when I reconfigure the FPGA.
I doubt there is a way to stop the system from restarting. I suggest checking the hardware log to see what is being detected and initiating the system shutdown. I suspect a critical event is occurring. If you are able to restart the PCIe card then it would be similar to hot-swapping the PCIe. It would go through link negotiation and firmware initialization again. I'm not aware of any of our servers that support PCIe hot-swap, but you can check the manual on the system support page.
I think I can't prevent the link to go down during reprogramming. Is it possible to prevent the bridge to which the FPGA is connected from forwarding this error to the root port?
I can't pretend to be an expert with FPGAs, but I'm trying to make one behave like a SEGA Genesis! On topic - my thought process is that perhaps the processor is reading the card in it's config as a "real" device, but when it gets reprogrammed, the processor thinks the card just fell offline and triggers a kernel panic. You might consider looking for software updates to whatever is reprogramming the FPGA, because maybe there is something that can communicate an "unmount" to the processor.
Good luck to you, I'll be interested to see the fix to this.
i am facing a similar issue on a Dell R620 but in other hand i am running a linux RouterOS system with 1 mezzanine 4 ports Intel slot + 1 dual port sfp+ 10G on top , whilst running trought put bandwith testing the router software troughput capacity i can reach using local 127.0.0.1 ip just for testing the 2 CPus E5-2697v2 i get get full throtle from the cpus running 88gbps full trougput TCP and 400gbps udp full troughput with cpus maxed out.. then i did some testing routing troughput from 1 server R620 to another server 630 via 2 mellanox mcx354 cards 40gbps cable after running the troughput test at around 55gbps full troughput being sent from one server to the other.. the R620 just reboots..
at first i tought it might be the pci-e card badly inserted in the slot.. , i shut down the server remove the pci-e 3.0 x8 mellanox card.. and inserted it back up.. the card starts, the server boots normaly with no errors at all i login in to the routeros and started testing again ant it reboots the server again showing up on the LCD system rebooting..
i tought it could be the16 GB memory modules.. so i swapped out the full 128GB ram i have 64GB ram running on each CPU slot memory module dimms.. replaced all of them with new one DDR3 1333mhz... but same error happened..
so i have decied to remove the mellanox mcx354 card and swap with another one "has we have several cards in stock" did new testing with new mellanox mcx354 card anda after good 20 seconds or so of troughput testing.. server reboots again... so i tought well the pci slot ix 3.0 x16 on the riser card so it can handle the full 64gbps troughput.. but instead i have replaced the network card with a 10G dual sfp+ card.. and on the server R630 same thing removed the other mellanox mcx354 and inserted a pair of mellanox dual sfp+ 10G on both servers...
fired up the 620 again and did a new test and to my suprise after it reached average 18gbs full troughput download/upload it reboots the R620.. the other server is ok no reboots..
once the server rebooted i have enabled a traffic generator server app with pppoe connections testing.. and with less speed at around 400mbps traffic troughput with high load of pppoe sessions.. the server reboots also system rebooting.. but on the LCD screen on errors shown.. i tought erros were disabled but they were not.. so i have replaced 1 power supply with server on.,. and the alarms came up on the LCD screen, removed RAID disk whilst on and alarm popped up..
i am actually running out of ideas.. could it be a faulty hardware? pci-e slot? riser card?
i dont think is the cpu has i have put both cpus to full throttle on cpu bandwitdh troughput and it goes nearly 95% withou rebooting i tested if for a good half-hour with server running on high rpm..
once i start inserting the pci-e slots to test network cards the issue starts..
BTW.. Dell bios updated latest version available on dell website, latest idrac, latest raid controler, latest power supply firmware patch, latest chipset drivers.. everything fully updated..
Hello, at this point my suspicion lies with faulty some riser fault. Is there anyway you can try a new one and see if you still face the same? Respectfully,
Social Media and Communities Professional
Dell Technologies | Enterprise Support Services
#IWork4Dell
Did I answer your query? Please click on ‘Mark as Accepted Answer’. ‘Thumbs up’ the posts you like!
Daniel My
12 Elder
•
6205 Posts
1912
0
Posted October 3rd, 2018 10:00
Hello
I doubt there is a way to stop the system from restarting. I suggest checking the hardware log to see what is being detected and initiating the system shutdown. I suspect a critical event is occurring. If you are able to restart the PCIe card then it would be similar to hot-swapping the PCIe. It would go through link negotiation and firmware initialization again. I'm not aware of any of our servers that support PCIe hot-swap, but you can check the manual on the system support page.
http://www.dell.com/support/
Thanks