I am running a Server 2008 R2 Hyper V cluster with 2 X R710 servers and a MD3000i.
Today suddenly my MD3000i decided it is time to play a little. This is the error generated - Event Message: Virtual Disk not on preferred path due to AVT/RDAC failover Component type: RAID Controller Module Component location: RAID Controller Module in slot 1
Looking at the whole cluster there is no network hardware failures and everything is on.
What I then did was to "Redistribute Virtual Disks" under "Manage RAID Controlllers" it then succeeds and the array is optimal for 5 seconds then the same error / issue occurs. According to the MDSM software both controllers are online. It does not make sense.
Next step I shutdown all the Virtual Machines running on the cluster ( not good for production environment ) . Suddenly when the last virtual machine was Offline the MD3000i changes its status to Optimal. I left it for 5 min and it stayed like that. I started to boot all the Virtual Machines, they are now all up and running as I type this. The SAN has not logged a error since.
What is happening here? Surely to stop and start all the Virtual machines could not solve the issue and all is running as it were before the error started. No it does not happen again !!!
I am running a Server 2008 R2 Hyper V cluster with 2 X R710 servers and a MD3000i.
Today suddenly my MD3000i decided it is time to play a little. This is the error generated - Event Message: Virtual Disk not on preferred path due to AVT/RDAC failover Component type: RAID Controller Module Component location: RAID Controller Module in slot 1
Looking at the whole cluster there is no network hardware failures and everything is on.
What I then did was to "Redistribute Virtual Disks" under "Manage RAID Controlllers" it then succeeds and the array is optimal for 5 seconds then the same error / issue occurs. According to the MDSM software both controllers are online. It does not make sense.
Next step I shutdown all the Virtual Machines running on the cluster ( not good for production environment ) . Suddenly when the last virtual machine was Offline the MD3000i changes its status to Optimal. I left it for 5 min and it stayed like that. I started to boot all the Virtual Machines, they are now all up and running as I type this. The SAN has not logged a error since.
What is happening here? Surely to stop and start all the Virtual machines could not solve the issue and all is running as it were before the error started. No it does not happen again !!!
Could anyone maybe assist ??
"
Are you running the Dell MPIO drivers? (fairly important), but even with that, Hyper V's failover and fail back sort of sucks. It seems like you just had a valid momentary loss of your path, then Hyper V was not able to switch it back on the fly.
I have experienced it with Hyper V and a decent number of various sans.
Yes I am running the Dell MPIO drivers.
Am I understanding you correctly that it is Hyper V causing the virtual disks on the SAN to failover to the second RAID controller ?
I don't think Hyper-V caused it. But once it happened, Hyper-V has no good failback to the correct / prefered path listed on the Virtual disk ownership in the MD3000i san. Thus Hyper-V is the one that is using the wrong path and the MD3000i is just reporting it as being wrong. So you fix it on the Md3000i and then Hyper-V just switches it back.
I think anyways. I am a little weak on Hyper-V, I'm still old school and still try my best to get people to pony up for ESX / VMware. We had a Hyper-V setup on our test bench for a long time on an MD3000i and I remember having similar issues, I also have trouble shot other Hyper-V installations with similar issues with other san models.
Do some googling on Hyper-V iSCSI san failover and failback. I think you may have to re-create a failure to get it to fail back correctly.
Clues to the solution should be found in the iSCSI initiator console on the Hyper-V host itself.
Hi Thanks for the reply, Ok I understand you now 100%. It has not happened again yet, but what did happen is that cluster reported that redirected access was turned on for one of the Cluster shared volumes. I had to shutdown all the virtual machines again and then the error dissipated. I think I need to get hold of good quality network cables and replace the current ones.
I checked the initiator and could not see any issues while in the crooked state.
I am under pressure to move more production servers over onto the cluster but in this state I do not want to.
JOHNADCO
2 Intern
•
847 Posts
444
0
Posted May 24th, 2010 08:00
I have experienced it with Hyper V and a decent number of various sans.