Unsolved
This post is more than 5 years old
1 Message
0
6535
May 6th, 2008 07:00
Powervault 220S Weird problem - Array mysteriously disappearing
Hi,
I'll give a brief history of our problem, we installed a poweredge 2950 with a powervault 220S at the end of 2006, it ran pretty much perfectly for 9 months and then the following happened.
I would come back to work and the server would show a "Windows Delayed Write Failure" having got through the error messages and looked at the server, our H Drive (logical drive for our array running on the 220S) had gone.
Phoned Dell (UK), shutdown the server and then the 220S, got through to Dell, we turned the 220S on waited for it to do it's post stuff and then booted up the 2950, checked the perc4 de bios and the array was fine. Booted into windows, MS Exchange 2003 ran it's recovery features and we were up and running. Dell thought it was an unseated cable and or controller on the 220S. They got me to power everything down and unseat and reseat everything. 3 months later and the problem surfaced again, Dell (UK) Tech, I think were not sure of what's going on, I'm not sure either, going through the logs the server seems to be doing it's normal thing and then all of a sudden we get the following :
Date: 03/05/2008
Time: 02:31:13
The device, \Device\Scsi\mraid35x1, did not respond with a timeout period.
Date 03/05/2008
Time: 02:34:13
The driver detected a controller error on \Device\Scsi\mraid35x1
Date 03/05/2008
Time: 02:34:19
Communication with the enclosure has been lost: Enclosure 1:6, Controller 0, Controller 1.
Then Windows has a coronary.
Dell (UK) thought it could of been a firmware issue so we updated the perc 4e to the lastest version December 2007. This brought 4 months of bliss, we thought we had finally resolved the problem to a dodgy firmware update. However at the end of April the issue has resurfaced. I ran DSET (again) and sent all the logs to Dell (UK), they went through the logs and they noticed psu 2 have a power fluctuation, but we don't have a 2nd psu in the 220s. We have redundant psu's in the 2950 but they're connected to alternate UPS systems to cover a ups failure.
Again Dell (UK) weren't sure of what's going on and said I would need to monitor it and come back if it happened again. I've contacted them again (email) and this time I thought I would post in the forum to see if anyone has also experienced this issue.
The times are at which point the raid controller loses contact with the array are random, the date at which it happens is fairly random. Prior to the fault, the events recorded in the eventlog are normal and then the above 3 are inserted and the array has gone.
Our 2950 only runs exchange 2003, nothing else.
It has 4 drives, 2 sets of mirrors.
0 & 1 are for OS and page file
3 & 4 are for the exchange transaction logs
The array on the 220S is a raid 10, it only holds the exchange database.
No other servers are connected to the 220S, can provide any more info if needed.
Kind regards
Lavan

