I have a test setup running 2.0.1.2. This is a 3 node Ubuntu linux16.04 setup with 2 networks, one management and the other data. The sds configuration is set for the data network (Communicate Both SDS and SDC). On one node I keep getting snmp traps MDM.MDM_Cluster.CLUSTER_DEGRADED and have found in the SDS trc.0 log events such as below that match the exact time of the trap.
Is there a guide or documentation on what Oscillation types are and what they mean or has anyone else had this one and know what is causing it? It is so brief, I haven't even see the admin interface display a problem and no rebuild occurs from what I see. It is truly is just a blip as can be seen in the 1 sec reference. In the Scaleio client under the SDS their are "none found" for Oscillating Failure Counters.
26/06 21:56:24.203890 0x7f0701bdaeb8:contNet_OscillationNotif:01720: Con ca6db64300000002 - Oscillation of type 5 (RPC_LINGERED_1SEC) reported
In most cases these errors indicate a "hiccup" in the network, causing temporary disconnections. Sometimes they can indicate disk problems (i.e. MDM running on a slow/faulting HDD) or CPU starvation (we mostly see it in virtualized environments though). I would check and tune both networks used by MDM, see if there are any dropped packets on the interfaces etc.
Osciliating errors are shortly described in ScaleIO Deployment Guide, I don't think there's any documentation that covers them in-depth, they are mostly for internal debugging.
I finally found in the MDM logs on a recent event this data. Below is a vew from MDM1 and MDM2 and TB1. Not really seeing much in the TB logs almost like it still sees MDM2 but MDM1 doesn't is what I surmise. I just want to focus on the problem machine/machines. The lost lease is also interesting and im not sure what the process is to determine lease.
10/07 10:38:50.268944 0x7fef80a10eb8:actor_Loop:11958: Still handling the same trigger since we couldn't convince all voters yet, votersFullyUpdated=0, voterHalf=1
10/07 10:38:50.270315 0x7fef80a10eb8:actorVoter_ProcessOneRsp:09502: We have the lease. Update degradedLocal: 83 Msg: 84
10/07 10:38:50.270319 0x7fef80a10eb8:actorVoter_ProcessOneRsp:09776: voterID: 6894583c50312f90 out of order response. bHasLease: 1 newLeaseTime: 2139617696 ticksExpiration: 2139617696
10/07 10:38:50.270322 0x7fef80a10eb8:actor_Loop:11958: Still handling the same trigger since we couldn't convince all voters yet, votersFullyUpdated=1, voterHalf=1
10/07 10:38:50.272084 0x7fef80a10eb8:actorVoter_ProcessOneRsp:09502: We have the lease. Update degradedLocal: 83 Msg: 84
10/07 10:38:50.272086 0x7fef80a10eb8:actorVoter_ProcessOneRsp:09776: voterID: 473f63f47cf794d2 out of order response. bHasLease: 1 newLeaseTime: 2139617696 ticksExpiration: 2139617696
10/07 10:38:50.272097 0x7fef80a10eb8:mosEventLog_PostInternal:00590: New event added. Message: "MDM cluster node is now DEGRADED - node ID 7b21663a21af5291; IPs: [SecondaryMDMIP], Port: 9011 is offline.". Additional info: "" Severity: Error
10/07 10:38:50.272101 0x7fef809e3eb8:syncerDegrador_Umt:01445: Degrador finished to move to state MOVE_TO_DEGRADED. bNextDegraded 1
10/07 10:38:50.272115 0x7fef80a19eb8:syncer_MoveToDegraded:00593: Syncer UMT finished moving to degraded mode
10/07 10:38:50.272125 0x7fef80a19eb8:syncer_SendStartSync:00480: syncSize: 624848. Local: PID 3845. Gen 626Msg: PID 3845. Gen 626
MDM2:
10/07 10:38:50.218350 0x7f8ad0a10eb8:actorLoop_NeedCede:11174: Not ceding as there are still enough free voters. voterHalf 1 ownedByOthers: [0,0] voteStates: [INVALID=0,UNKNOWN=3,NOT_OWNED=0,OWNED_BY_ME=0,OWNED_BY_OTHER=0,BLOCKED=0,NO_ANSWER=0,ERROR=0]
10/07 10:38:50.218393 0x7f8ad0a10eb8:actorLoop_NeedCede:11174: Not ceding as there are still enough free voters. voterHalf 1 ownedByOthers: [1,0] voteStates: [INVALID=0,UNKNOWN=2,NOT_OWNED=0,OWNED_BY_ME=0,OWNED_BY_OTHER=1,BLOCKED=0,NO_ANSWER=0,ERROR=0]
10/07 10:38:50.218435 0x7f8ad0a10eb8:actorLoop_NeedCede:11174: Not ceding as there are still enough free voters. voterHalf 1 ownedByOthers: [1,0] voteStates: [INVALID=0,UNKNOWN=2,NOT_OWNED=0,OWNED_BY_ME=0,OWNED_BY_OTHER=1,BLOCKED=0,NO_ANSWER=0,ERROR=0]
10/07 10:38:50.218455 0x7f8ad0a10eb8:actorLoop_NeedCede:11174: Not ceding as there are still enough free voters. voterHalf 1 ownedByOthers: [0,0] voteStates: [INVALID=0,UNKNOWN=3,NOT_OWNED=0,OWNED_BY_ME=0,OWNED_BY_OTHER=0,BLOCKED=0,NO_ANSWER=0,ERROR=0]
10/07 10:38:50.319026 0x7f8ad0a10eb8:actorLoop_NeedCede:11174: Not ceding as there are still enough free voters. voterHalf 1 ownedByOthers: [0,0] voteStates: [INVALID=0,UNKNOWN=3,NOT_OWNED=0,OWNED_BY_ME=0,OWNED_BY_OTHER=0,BLOCKED=0,NO_ANSWER=0,ERROR=0]
10/07 10:38:50.319199 0x7f8ad0a10eb8:actor_Loop:11775: Set new degraded gen 84,old one was. 83.New degraded [25b3bc8d1145f4f1] old one []
10/07 10:38:50.319203 0x7f8ad0a10eb8:actorLoop_NeedCede:11174: Not ceding as there are still enough free voters. voterHalf 1 ownedByOthers: [1,0] voteStates: [INVALID=0,UNKNOWN=2,NOT_OWNED=0,OWNED_BY_ME=0,OWNED_BY_OTHER=1,BLOCKED=0,NO_ANSWER=0,ERROR=0]
10/07 10:38:50.323629 0x7f8ad0a07eb8:voter_ReleaseMaster:01534: Releasing master - no successor, RC: SUCCESS
pawelw1
306 Posts
1828
0
Posted July 4th, 2017 02:00
Hi,
In most cases these errors indicate a "hiccup" in the network, causing temporary disconnections. Sometimes they can indicate disk problems (i.e. MDM running on a slow/faulting HDD) or CPU starvation (we mostly see it in virtualized environments though). I would check and tune both networks used by MDM, see if there are any dropped packets on the interfaces etc.
Osciliating errors are shortly described in ScaleIO Deployment Guide, I don't think there's any documentation that covers them in-depth, they are mostly for internal debugging.
Hope that helps!
Cheers,
Pawel