
UNSOLVED
Isilon Platform API service issue
We have an issue on our cluster (Currently running OneFS 7.2.0.1) where the Platform API service (papi) has started to do some wierd things. I started to receive errors when commands were run on the command line such as "incomplete response from server". This happened on 8 of the 13 nodes in the cluster wihch was a bit odd. The only difference between the 8 i got the error on and the 5 that were fine is these 5 are currently suspended in the IP Pools. At the same time we see this we see issues with getting into the WebUI and also InsightIQ connection to the cluster.
I've read the article about isi_papi_d being in a bad state and if i restart the service it works for about 30-40 minutes then starts to fail again.
I noticed that when it started to fail the process was always at the top of the list if i ran a top command and usually chewing up quite a bit of CPU% and either in a ucond stater or kqread state. And from what i can find neither of those is good as it seems to indicate that the process is waiting on something, somewhere.
I've tried the usual Windows fix or rebooting all the nodes since some had been up for well over 200 days and thats not worked.
So i thought i would ask the community, see if anyone has any thoughts on what could be causing this and how to get round it? All thoughts gratefully received.
Responses (0)
Solutions (0)
