NMDA: DB2-Backups schlagen jede Nacht zufällig mit Fehler 3 fehl
Summary: NMDA DB2-Backups schlugen letzte Nacht mit Fehler 3 fehl. Das Problem wurde nach der Erstellung eines neuen Geräts, der Verteilung von Backups auf zwei Speicher-Nodes und der Einrichtung von DB2-Wiederholungs- und Timeout-Parametern behoben. ...
Symptoms
NMDA-DB2-Backup schlägt fehl mit Error 3
DB2-Backup schlägt fehl mit Fehler 'lgto_auth for `nsrmmd' failed: busy"
Es wurden keine Netzwerk- oder Firewallprobleme gefunden.
Es gibt 1000e der folgenden Meldungen in /nsr/logs/daemon.raw Im Storage Node:
"5004-nfs lookup failed (nfs: No such file or directory)""invalid save stream""Cannot stat active file""unable to collect deduplication statistics""was aborted and removed from volume"
Fehler in nmda-messages.log libnsrdb2.log möglicherweise nicht mit debug=9verwalten:
153929 2/9/2021 10:34:50 PM 4 7 987 1 18153790 0 (client) (pid18153790) NSR severe The backup session could not start: busy.
93412 2/9/2021 10:34:50 PM 3 5 0 1 18153790 0 (client) (pid18153790) NSR error Could not perform the action 2. The status was changed to 3.
153929 1612842069 4 7 987 1 19136950 0 (client) (pid19136950) NSR severe 39 The backup session could not start: %s. 1 49 8 0 4 busy
93412 1612842069 3 5 0 1 19136950 0 (client) (pid19136950) NSR error 62 Could not perform the action %d. The status was changed to %d. 2 1 1 2 1 1 3
(pid = 18809144) (02/09/21 21:40:00.338942) nsrdb2sv_log_program_args: /usr/bin/nsrdasv -LL -T db2 -s (NW server) -g (group) -a *policy action jobid=2297950 -a *policy name=(policy) -a *policy workflow name=(workflow) -a *policy action name=(action) -y Tue Feb 23 23:59:59 GMT-0600 2021 -w Tue Feb 23 23:59:59 GMT-0600 2021 -m (client) -a *policy action jobid restart=Yes -b (pool) -t 1612810625 -o ....
(pid = 18809144) (02/09/21 21:40:00.624767) Backing up the (DB) database.
(pid = 18809144) (02/09/21 21:40:00.624939) set_db2_version: Exiting set_db2_version(): Return code: 10050000
(pid = 18809144) (02/09/21 21:49:08.731480) DbBackup: Exiting with error:
Unable to backup DB2MDME database due to backup request failure, SQLCODE : -2025, SQL2025N An I/O error occurred. Error code: "3". Media on which this error occurred: "VENDOR".
.
(pid = 18809144) (02/09/21 21:49:08.731631) libdb2sv_main: ERROR: DbBackup() failed.
(pid = 18809144) (02/09/21 21:49:08.731685) Unable to backup DB2MDME database due to backup request failure, SQLCODE : -2025, SQL2025N An I/O error occurred. Error code: "3". Media on which this error occurred: "VENDOR".
Kritischer Fehler: nsrmmd Auslastungsfehler unten:
02/09/21 21:32:46 (pid 18153790): 02/09/21 21:32:46.797073 lgto_auth for `nsrd' succeeded 02/09/21 21:32:46 (pid 18153790): 02/09/21 21:32:46.855631 lgto_parms for `nsrmmd' succeeded 02/09/21 21:32:46 (pid 18153790): 02/09/21 21:32:46.855705 got `store index entries' value of `Yes' 02/09/21 21:32:46 (pid 18153790): 02/09/21 21:32:46.855803 Saving in pool 'IDC-DB2'. 02/09/21 21:32:46 (pid 18153790): 02/09/21 21:32:46.855822 server enabled for immediate mode 02/09/21 21:32:46 (pid 18153790): 02/09/21 21:32:46.882267 lgto_auth for `nsrmmd' failed: busy 02/09/21 21:32:46 (pid 18153790): 02/09/21 21:32:46.882349 Unable to acquire the user credentials for direct save nsrmmd authentication: busy. 02/09/21 21:32:46 (pid 18153790): 02/09/21 21:32:46.882439 The error TYPE is 0, SEVERITY is 0, NUMBER is -13, errnum is -13, errstr is 'busy'.
Cause
Resolution
Das Problem wurde nach den folgenden Änderungen behoben. Es gibt nicht die eine Grundursache, aber das Erstellen eines neuen Geräts und das Festlegen der folgenden Parameter hat am meisten geholfen:
1. Ein neues Gerät wurde zum Speicher-Node hinzugefügt.
2. Gleichmäßig auf die Storage Nodes verteilte Backups (Zielsitzung).
3. Geänderte Backupstartzeiten.
4. Die folgenden Parameter wurden in NMDA DB2-Anwendungsinformationen hinzugefügt:
NSR_MAX_START_RETRIES=50
NSR_FXBUSY_RETRIES=10
NSR_MMDB_RETRY_TIME=10
5. Erhöhung des Inaktivitäts-Timeouts auf 300, Wiederholungen=2, Wiederholungsverzögerung=10 in den Eigenschaften der Backupaktion.