PowerFlex MDM-klynge nede efter gentagne fejl
Summary: MDM-klyngen mister synkronisering gentagne gange og bliver til sidst utilgængelig / forbliver nede, indtil brugerne griber ind.
Symptoms
Beskrivelse af problem
MDM-klyngen mister synkronisering gentagne gange og bliver til sidst utilgængelig / forbliver nede, indtil brugerne griber ind.
Scenario
MDM-processen genstarter for hurtigt, og efter tid forhindrer systemd flere starter.
Hvis ingen synkroniserede sekundære MDM er kan påtage sig den primære rolle, vil systemet ikke have nogen primær MDM.
Symptomer
I MDM-hændelsesloggen er der mange tab af MDM-klyngeforbindelser:
2020-12-03 17:40:36.068 REMOTE_SYSLOG_MODULE_INITIALIZED INFO Initialized the remote syslog module 2020-12-03 17:40:36.068 MDM_MANAGER_START INFO MDM started with the role of Manager 2020-12-03 17:40:36.251 MDM_CLUSTER_CONNECTED INFO The MDM, MDM2 (ID 0f8dcd34388a8c01), connected after 0ms 2020-12-03 17:40:36.251 MDM_CLUSTER_CONNECTED INFO The MDM, MDM3 (ID 551526045129a502), connected after 0ms 2020-12-03 17:40:36.251 MDM_CLUSTER_CONNECTED INFO The MDM, TB2 (ID 7d5c0c020abdea04), connected after 0ms 2020-12-03 17:40:36.251 MDM_CLUSTER_CONNECTED INFO The MDM, TB1 (ID 4207e70f08980503), connected after 0ms 2020-12-03 17:40:37.486 MDM_CLUSTER_LOST_CONNECTION WARNING The MDM, MDM3 (ID 551526045129a502), has lost connection to the cluster. 2020-12-03 17:40:37.785 MDM_CLUSTER_CONNECTED INFO The MDM, MDM3 (ID 551526045129a502), connected after 310ms 2020-12-03 17:40:38.755 MDM_CLUSTER_LOST_CONNECTION WARNING The MDM, MDM3 (ID 551526045129a502), has lost connection to the cluster. 2020-12-03 17:40:39.060 MDM_CLUSTER_CONNECTED INFO The MDM, MDM3 (ID 551526045129a502), connected after 310ms 2020-12-03 17:40:40.032 MDM_CLUSTER_LOST_CONNECTION WARNING The MDM, MDM3 (ID 551526045129a502), has lost connection to the cluster. 2020-12-03 17:40:40.337 MDM_CLUSTER_CONNECTED INFO The MDM, MDM3 (ID 551526045129a502), connected after 310ms 2020-12-03 17:40:41.364 MDM_CLUSTER_LOST_CONNECTION WARNING The MDM, MDM3 (ID 551526045129a502), has lost connection to the cluster. 2020-12-03 17:40:41.673 MDM_CLUSTER_CONNECTED INFO The MDM, MDM3 (ID 551526045129a502), connected after 310ms 2020-12-03 17:40:42.602 MDM_CLUSTER_NOT_RESPOND WARNING The MDM, MDM1 (ID 30deabdb5ddf2a00), is not responding 2020-12-03 17:40:42.676 MDM_CLUSTER_LOST_CONNECTION WARNING The MDM, MDM3 (ID 551526045129a502), has lost connection to the cluster. 2020-12-03 17:40:43.091 MDM_CLUSTER_CONNECTED INFO The MDM, MDM3 (ID 551526045129a502), connected after 410ms 2020-12-03 17:40:44.073 MDM_CLUSTER_LOST_CONNECTION WARNING The MDM, MDM3 (ID 551526045129a502), has lost connection to the cluster. 2020-12-03 17:40:45.557 MDM_CLUSTER_LOST_CONNECTION WARNING The MDM, MDM2 (ID 0f8dcd34388a8c01), has lost connection to the cluster. 2020-12-03 17:40:45.858 MDM_CLUSTER_CONNECTED INFO The MDM, MDM2 (ID 0f8dcd34388a8c01), connected after 310ms 2020-12-03 17:40:46.967 MDM_CLUSTER_LOST_CONNECTION WARNING The MDM, MDM2 (ID 0f8dcd34388a8c01), has lost connection to the cluster. 2020-12-03 17:40:47.268 MDM_CLUSTER_CONNECTED INFO The MDM, MDM2 (ID 0f8dcd34388a8c01), connected after 300ms 2020-12-03 17:40:48.413 MDM_CLUSTER_LOST_CONNECTION WARNING The MDM, MDM2 (ID 0f8dcd34388a8c01), has lost connection to the cluster. 2020-12-03 17:40:48.625 MDM_CLUSTER_NOT_RESPOND WARNING The MDM, MDM3 (ID 551526045129a502), is not responding 2020-12-03 17:40:48.811 MDM_CLUSTER_CONNECTED INFO The MDM, MDM2 (ID 0f8dcd34388a8c01), connected after 400ms 2020-12-03 17:40:49.866 MDM_CLUSTER_LOST_CONNECTION WARNING The MDM, MDM2 (ID 0f8dcd34388a8c01), has lost connection to the cluster. 2020-12-03 17:40:50.160 MDM_CLUSTER_CONNECTED INFO The MDM, MDM2 (ID 0f8dcd34388a8c01), connected after 300ms 2020-12-03 17:40:51.208 MDM_CLUSTER_LOST_CONNECTION WARNING The MDM, MDM2 (ID 0f8dcd34388a8c01), has lost connection to the cluster. 2020-12-03 17:40:51.520 MDM_CLUSTER_CONNECTED INFO The MDM, MDM2 (ID 0f8dcd34388a8c01), connected after 310ms 2020-12-03 17:40:52.603 MDM_CLUSTER_LOST_CONNECTION WARNING The MDM, MDM2 (ID 0f8dcd34388a8c01), has lost connection to the cluster. 2020-12-03 17:40:53.407 MDM_CLUSTER_BECOMING_MASTER WARNING This MDM, MDM1 (ID 30deabdb5ddf2a00), took control of the cluster and is now the Master MDM. 2020-12-03 17:40:53.619 MDM_CLUSTER_NODE_DEGRADED ERROR MDM cluster node is now DEGRADED and is in offline node MDM3 (ID 551526045129a502); IPs: [10.180.88.5], Port: 9011 . 2020-12-03 17:40:53.619 MDM_CLUSTER_NODE_DEGRADED ERROR MDM cluster node is now DEGRADED and is in offline node MDM2 (ID 0f8dcd34388a8c01); IPs: [10.180.88.69], Port: 9011 . 2020-12-03 17:40:54.622 REMOTE_SYSLOG_MODULE_INITIALIZED INFO Initialized the remote syslog module 2020-12-03 17:40:54.622 MDM_MANAGER_START INFO MDM started with the role of Manager 2020-12-03 17:40:54.743 MDM_CLUSTER_CONNECTED INFO The MDM, TB2 (ID 7d5c0c020abdea04), connected after 0ms 2020-12-03 17:40:54.842 MDM_CLUSTER_CONNECTED INFO The MDM, TB1 (ID 4207e70f08980503), connected after 0ms 2020-12-03 17:40:54.943 MDM_CLUSTER_BECOMING_MASTER WARNING This MDM, MDM1 (ID 30deabdb5ddf2a00), took control of the cluster and is now the Master MDM. 2020-12-03 17:40:55.144 MDM_CLUSTER_NODE_DEGRADED ERROR MDM cluster node is now DEGRADED and is in offline node MDM3 (ID 551526045129a502); IPs: [10.180.88.5], Port: 9011 . 2020-12-03 17:40:55.145 MDM_CLUSTER_NODE_DEGRADED ERROR MDM cluster node is now DEGRADED and is in offline node MDM2 (ID 0f8dcd34388a8c01); IPs: [10.180.88.69], Port: 9011 . 2020-12-03 17:40:56.140 REMOTE_SYSLOG_MODULE_INITIALIZED INFO Initialized the remote syslog module 2020-12-03 17:40:56.140 MDM_MANAGER_START INFO MDM started with the role of Manager 2020-12-03 17:40:56.229 MDM_CLUSTER_CONNECTED INFO The MDM, TB2 (ID 7d5c0c020abdea04), connected after 0ms 2020-12-03 17:40:56.327 MDM_CLUSTER_CONNECTED INFO The MDM, TB1 (ID 4207e70f08980503), connected after 0ms 2020-12-03 17:40:56.428 MDM_CLUSTER_BECOMING_MASTER WARNING This MDM, MDM1 (ID 30deabdb5ddf2a00), took control of the cluster and is now the Master MDM. 2020-12-03 17:40:56.629 MDM_CLUSTER_NODE_DEGRADED ERROR MDM cluster node is now DEGRADED and is in offline node MDM3 (ID 551526045129a502); IPs: [10.180.88.5], Port: 9011 . 2020-12-03 17:40:56.629 MDM_CLUSTER_NODE_DEGRADED ERROR MDM cluster node is now DEGRADED and is in offline node MDM2 (ID 0f8dcd34388a8c01); IPs: [10.180.88.69], Port: 9011 . 2020-12-03 17:40:57.660 REMOTE_SYSLOG_MODULE_INITIALIZED INFO Initialized the remote syslog module 2020-12-03 17:40:57.660 MDM_MANAGER_START INFO MDM started with the role of Manager 2020-12-03 17:40:57.768 MDM_CLUSTER_CONNECTED INFO The MDM, TB2 (ID 7d5c0c020abdea04), connected after 0ms 2020-12-03 17:40:57.869 MDM_CLUSTER_CONNECTED INFO The MDM, TB1 (ID 4207e70f08980503), connected after 0ms 2020-12-03 17:40:57.970 MDM_CLUSTER_BECOMING_MASTER WARNING This MDM, MDM1 (ID 30deabdb5ddf2a00), took control of the cluster and is now the Master MDM. 2020-12-03 17:40:58.171 MDM_CLUSTER_NODE_DEGRADED ERROR MDM cluster node is now DEGRADED and is in offline node MDM3 (ID 551526045129a502); IPs: [10.180.88.5], Port: 9011 . 2020-12-03 17:40:58.172 MDM_CLUSTER_NODE_DEGRADED ERROR MDM cluster node is now DEGRADED and is in offline node MDM2 (ID 0f8dcd34388a8c01); IPs: [10.180.88.69], Port: 9011 . 2020-12-03 17:40:59.143 REMOTE_SYSLOG_MODULE_INITIALIZED INFO Initialized the remote syslog module 2020-12-03 17:40:59.144 MDM_MANAGER_START INFO MDM started with the role of Manager 2020-12-03 17:40:59.245 MDM_CLUSTER_CONNECTED INFO The MDM, TB2 (ID 7d5c0c020abdea04), connected after 0ms 2020-12-03 17:40:59.353 MDM_CLUSTER_CONNECTED INFO The MDM, TB1 (ID 4207e70f08980503), connected after 0ms 2020-12-03 17:40:59.454 MDM_CLUSTER_BECOMING_MASTER WARNING This MDM, MDM1 (ID 30deabdb5ddf2a00), took control of the cluster and is now the Master MDM. 2020-12-03 17:40:59.655 MDM_CLUSTER_NODE_DEGRADED ERROR MDM cluster node is now DEGRADED and is in offline node MDM3 (ID 551526045129a502); IPs: [10.180.88.5], Port: 9011 . 2020-12-03 17:40:59.655 MDM_CLUSTER_NODE_DEGRADED ERROR MDM cluster node is now DEGRADED and is in offline node MDM2 (ID 0f8dcd34388a8c01); IPs: [10.180.88.69], Port: 9011 . 2020-12-03 17:41:00.630 REMOTE_SYSLOG_MODULE_INITIALIZED INFO Initialized the remote syslog module 2020-12-03 17:41:00.630 MDM_MANAGER_START INFO MDM started with the role of Manager 2020-12-03 17:41:00.722 MDM_CLUSTER_CONNECTED INFO The MDM, TB2 (ID 7d5c0c020abdea04), connected after 0ms 2020-12-03 17:41:00.818 MDM_CLUSTER_CONNECTED INFO The MDM, TB1 (ID 4207e70f08980503), connected after 0ms 2020-12-03 17:41:00.919 MDM_CLUSTER_BECOMING_MASTER WARNING This MDM, MDM1 (ID 30deabdb5ddf2a00), took control of the cluster and is now the Master MDM. 2020-12-03 17:41:01.120 MDM_CLUSTER_NODE_DEGRADED ERROR MDM cluster node is now DEGRADED and is in offline node MDM3 (ID 551526045129a502); IPs: [10.180.88.5], Port: 9011 . 2020-12-03 17:41:01.121 MDM_CLUSTER_NODE_DEGRADED ERROR MDM cluster node is now DEGRADED and is in offline node MDM2 (ID 0f8dcd34388a8c01); IPs: [10.180.88.69], Port: 9011 . 2020-12-03 20:37:38.973 REMOTE_SYSLOG_MODULE_INITIALIZED INFO Initialized the remote syslog module 2020-12-03 20:37:38.973 MDM_MANAGER_START INFO MDM started with the role of Manager
Hvilket betyder, at MDM-processer genstarter mange gange i træk. Dette kan skyldes alvorlige forbindelsesproblemer, eller at OS-disken reagerer langsomt som i Langsom skrivninger til OS-disk kan forårsage flere MDM-problemer. ("Hærden tog for lang tid") I journalctl eller /var/log/messages: "Started scaleio mdm" vises umiddelbart efter, at tjenesten stopper, hvis systemd genstarter mdm-tjenesten. Når systemd ikke omlægger det til at genstarte, ser du "gentaget for hurtigt"/"kunne ikke starte" (som kl. 17:41:01 nedenfor):
Dec 3 17:40:54 RHEL7-1 systemd: mdm.service: main process exited, code=exited, status=255/n/a Dec 3 17:40:54 RHEL7-1 systemd: Unit mdm.service entered failed state. Dec 3 17:40:54 RHEL7-1 systemd: mdm.service failed. Dec 3 17:40:54 RHEL7-1 systemd: mdm.service has no holdoff time, scheduling restart. Dec 3 17:40:54 RHEL7-1 systemd: Stopped scaleio mdm. Dec 3 17:40:54 RHEL7-1 systemd: Started scaleio mdm. Dec 3 17:40:55 RHEL7-1 systemd: mdm.service: main process exited, code=exited, status=255/n/a Dec 3 17:40:55 RHEL7-1 systemd: Unit mdm.service entered failed state. Dec 3 17:40:55 RHEL7-1 systemd: mdm.service failed. Dec 3 17:40:55 RHEL7-1 systemd: mdm.service has no holdoff time, scheduling restart. Dec 3 17:40:55 RHEL7-1 systemd: Stopped scaleio mdm. Dec 3 17:40:55 RHEL7-1 systemd: Started scaleio mdm. Dec 3 17:40:57 RHEL7-1 systemd: mdm.service: main process exited, code=exited, status=255/n/a Dec 3 17:40:57 RHEL7-1 systemd: Unit mdm.service entered failed state. Dec 3 17:40:57 RHEL7-1 systemd: mdm.service failed. Dec 3 17:40:57 RHEL7-1 systemd: mdm.service has no holdoff time, scheduling restart. Dec 3 17:40:57 RHEL7-1 systemd: Stopped scaleio mdm. Dec 3 17:40:57 RHEL7-1 systemd: Started scaleio mdm. Dec 3 17:40:58 RHEL7-1 systemd: mdm.service: main process exited, code=exited, status=255/n/a Dec 3 17:40:58 RHEL7-1 systemd: Unit mdm.service entered failed state. Dec 3 17:40:58 RHEL7-1 systemd: mdm.service failed. Dec 3 17:40:58 RHEL7-1 systemd: mdm.service has no holdoff time, scheduling restart. Dec 3 17:40:58 RHEL7-1 systemd: Stopped scaleio mdm. Dec 3 17:40:58 RHEL7-1 systemd: Started scaleio mdm. Dec 3 17:41:00 RHEL7-1 systemd: mdm.service: main process exited, code=exited, status=255/n/a Dec 3 17:41:00 RHEL7-1 systemd: Unit mdm.service entered failed state. Dec 3 17:41:00 RHEL7-1 systemd: mdm.service failed. Dec 3 17:41:00 RHEL7-1 systemd: mdm.service has no holdoff time, scheduling restart. Dec 3 17:41:00 RHEL7-1 systemd: Stopped scaleio mdm. Dec 3 17:41:00 RHEL7-1 systemd: Started scaleio mdm. Dec 3 17:41:01 RHEL7-1 systemd: mdm.service: main process exited, code=exited, status=255/n/a Dec 3 17:41:01 RHEL7-1 systemd: Unit mdm.service entered failed state. Dec 3 17:41:01 RHEL7-1 systemd: mdm.service failed. Dec 3 17:41:01 RHEL7-1 systemd: mdm.service has no holdoff time, scheduling restart. Dec 3 17:41:01 RHEL7-1 systemd: Stopped scaleio mdm. Dec 3 17:41:01 RHEL7-1 systemd: start request repeated too quickly for mdm.service Dec 3 17:41:01 RHEL7-1 systemd: Failed to start scaleio mdm. Dec 3 17:41:01 RHEL7-1 systemd: Unit mdm.service entered failed state. Dec 3 17:41:01 RHEL7-1 systemd: mdm.service failed.
Kort sagt kan du finde, hvornår systemd stoppede med at bringe MDM-tjenesten op igen med følgende linjer:
systemd: start request repeated too quickly for mdm.service systemd: Failed to start scaleio mdm.
Effektdata er ikke tilgængelige.
Cause
En del af MDM's adfærd, når den giver den primære rolle, er at genstarte (planlagt nedbrud) processen.
Systemd in RHEL/CentOS 6.x & 7.x har en tærskel for, hvor mange gange en proces kan genstarte inden for en bestemt tidsramme.
Hvis MDM-tjenesten ikke reagerer gentagne gange nok, tillader systemd ikke, at den genstarter.
Resolution
Løsning
- Stabiliser MDM-klyngen.
- Hvis det ikke er muligt, skal du kontakte SIO L3 for at få hjælp og citere denne artikel.
Påvirkede versioner
RHEL/CentOS 6.x og 7.x
Løst i version
TBD