Unsolved
This post is more than 5 years old
17 Posts
0
15981
November 9th, 2014 23:00
MD3200 dual controller SAS disconnecting single physical host - need help please
hi guys
We have 3 X R720 with SAS HBA (dual port) connected to an MD3200 (dual controller). It is setup in duplex mode and so should allow failover between the controllers. However on random occasions and random hosts it disconnects completely and the host has to be rebooted to enable connection to the storage.
All 3 hosts are VMWARE. Last week we migrated everything to a MD1200 (single controller) and for a week no drops.
On the weekend we migrated back to the MD3200 and it dropped this morning.
MD1200 - 1 disk group - 4 LUN's - default group has all LUN's included
MD3200 - same config as above.
Alternate LUN's set to alternate controllers as owner
I have the MD logs and ESX logs.
Has anyone seen this very strange behavior before?
Thanks


DELL-Sam L
Community Manager
•
8K Posts
•
357 Points
0
November 10th, 2014 13:00
Hello SeanBezuids,
Are all 3xR720 in a cluster or no? You stated that you have 1 disk group & 4 luns in the group, do all the host have access to all 4 luns or does each host have just access to 1 lun? Is it the same lun that disconnects from the different host or is it a different lun every time?
Please let us know if you have any other questions.
SeanBezuids
17 Posts
1
November 11th, 2014 00:00
Hi Sam
The 3 hosts are not clustered but all running VMware 5.5 U2.
Each host has a dual port HBA 6GB SAS with dual ports. Each port has a connection into alternate controllers on the MD.
Symptoms: Random disconnects of hosts. No fixed time, no errors in VMWare logs, HBA card goes offline completely. Only a reboot restores connectivity.
No errors reported in MD logs. The other hosts on the same controller continue to function without issue.
We implemented the following changes last night:
1. HBA driver version in ESX. There is an updated driver published after 5.5U1 came out. The old is version 14 and the new version 19.
2. ESX setting called “Interrupt remapping” - This is enabled by default but can cause many of the symptoms we have seen.
The Dell VMware ISO has interrupt mapping disabled by default. This was used to build the hosts however it was enabled. we are unsure how this changed however it could be the cause of the issue.
**PLEASE NOTE** The error message pointing to the interrupt mapping issue has the word APIC in it. I will find the right message. However this will not show in VMWARE logs if it boots off local storage. This is why we were unable to pick it up at all.
We updated the hosts last night and then tested by removing one of the SAS cables with VM's running. The failover to the other controller was seamless as expected.
I will update this in a week's time if there are no further issues.