NetWorker:復原失敗,並顯示「NDMP 服務錯誤:無法識別格式」
Summary: 由於網路資料管理通訊協定 (NDMP) 標頭損毀,導致 I/O 錯誤和檔案編號中斷定位,因此還原失敗,觸發「NDMP 服務錯誤:無法識別格式。」
Symptoms
這些症狀與兩個主要問題有關:
- 備份系統管理員未密切注意持續重複的 NMDP 標頭 I/O 錯誤,此 NetWorker 警告帶有 VNX 錯誤,應認真處理
- 兩種產品 (NetWorker 和 VNX) 將 NMDP 標頭損毀視為微不足道的錯誤。備份重啟的特殊情況:損壞達到嘗試讀取數據時無法自動解決的點
恢復所需的檔跨越多個備份(完整 + 差異 + 增量)並跨越多個卷。
NDMP 儲存集復原和逐個檔案 (可瀏覽) 復原失敗,並出現密切相關的錯誤
可瀏覽復原失敗記錄:
42744:nsrndmp_recover: Tape server paused: reached the end of file
42897:nsrndmp_recover: Opened the tape device : c208t0l0
42870:nsrndmp_recover: Continuing recover from the next volume
42619:nsrndmp_recover: NDMP Service Error: Cannot identify format. <<<<<<<<<<<<<<<<<<<<<<<<<<<<<<<<<<<<<<<<<<<<<<<<<<<<<<
42617:nsrndmp_recover: NDMP Service Log: server_archive: emctar vol 1, 0 files, 61440 bytes read, 0 bytes written
42738:nsrndmp_recover: Data server halted: Error during the restore.
42856:nsrndmp_recover: NDMP data server has an internal error.
42871:nsrndmp_recover: Error during File NDMP Extraction.
42840:nsrndmp_recover: NDMP recover failed.
42880:nsrndmp_recover: Error during NDMP recover
16279:winworkr: NDMP retrieval: child failed with status of 1
42897:nsrndmp_recover: Opened the tape device : c208t0l5
Mover listen address is NULL(NDMP_ADDR_LOCAL)
只有完整備份會發生儲存集復原失敗,其他層級的備份可使用儲存集復原來復原。
root@NW-server# nsrndmp_recover -s NW-server -c NAS-host -S 1011119015 -v off -m "NAS-host::/root_vdm_1/users/recover_emc" "/root_vdm_1/users/VDI_Users/v009/project/Plan04.2012"
42879:nsrndmp_recover: Peforming recover with no file mark dependency..
05/31/16 17:25:50.504040 NDMP Service Debug: The process id for NDMP service is 0xc4bd90b0
42787:nsrndmp_recover: Performing recover from NDMP type of device
05/31/16 17:26:01.269862 NDMP Service Debug: The process id for NDMP service is 0xc4bd90b0
May 31 17:26:02 NW-server root: [ID 702911 local0.alert] NetWorker media: (waiting) waiting for LTO Ultrium-5 tape VN0116L5 on NAS-host
95555:nsrmmd: ndmp tape mtio failed, I/O error
42850:nsrndmp_recover: Failed to load volume 1111782299 fnum 3 <<<<<< cannot load File Number 3
42855:nsrndmp_recover: Failed to load the tape.
42871:nsrndmp_recover: Error during File NDMP Extraction.
42840:nsrndmp_recover: NDMP recover failed.
42880:nsrndmp_recover: Error during NDMP recover
root@NW-server #
最重要的癥狀是掃描程式無法流覽所需的備份卷,它會在讀取任何數據之前過早停止!
root@NW-server # scanner -vvv -p "rd=NAS-host:c208t0l5 (NDMP)"
8909:scanner: using 'rd=NAS-host:c208t0l5 (NDMP)' as the device name
05/31/16 17:29:08.919423 NDMP Service Debug: The process id for NDMP service is 0xda018d60
9040:scanner: Opened c208t0l5 for read
8968:scanner: Reading the label...
8969:scanner: Reading the label done
8936:scanner: scanning LTO Ultrium-5 tape VN0116L5 on rd=NAS-host:c208t0l5 (NDMP)
96367:scanner: volume id 1111782299 record size 262144 bytes
created 5/24/16 21:00:35 expires 5/24/18 21:00:35
8973:scanner: setting position from fn 0, rn 0 to fn 2, rn 0
05/31/16 17:29:09.073276 NDMP Service Debug: The process id for NDMP service is 0xda018d60
9040:scanner: Opened c208t0l5 for read
05/31/16 17:29:09.201476 NDMP Service Debug: The process id for NDMP service is 0xda018d60
9040:scanner: Opened c208t0l5 for read
8761:scanner: done with LTO Ultrium-5 tape VN0116L5 <<<<<<<< No Errors , no mention of any file after File number 2 !!
NW 除錯標誌檔案的用法: /nsr/debu/ndmp_auto_pos 無法協助完整備份儲存集或可瀏覽復原。
Cause
NDMP 標頭損毀會導致格式錯誤的 NDMP 磁碟區格式,有時會產生錯誤的檔案編號 (例如自動備份重新開機):
檢查 NW 記錄顯示,寫入 NDMP 標頭時,每個備份都會伴有 I/O 錯誤。這是每次備份都會發生的一致問題,這是指向持續性損壞的指標。
70896 05/01/16 12:18:05 0 0 2 1 23177 0 NW-server nsrd NSR info NAS-host:/root_vdm_1/userdata saving to pool 'NDMPPool' (VN0050L5)
42597 05/01/16 12:18:05 2 0 0 1 23990 0 NW-server nsrmmd NSR warning ndmp header: I/O error <<<<<<<<<<<<<<<<<<<<<<<<<<<<<<<<<<<<<<<<<<<<<<<<<<
71193 05/01/16 12:18:05 0 0 0 1 23177 0 NW-server nsrd NSR info NDMP Save Notice: NAS-host:/root_vdm_1/users NDMP save running on 'NW-server'
檢查 NAS 記錄會顯示 NDMP 裝置的緩衝大小設定為小於在 NW 中設定的裝置區塊大小的值。預設的 NetWorker 組態會將 LTO-5 裝置的裝置區塊大小設定為 256 K,但 NDMP 裝置緩衝大小 (NDMP.bufsz) 為 128 K。此部分僅與 NDMP 標頭相關,因為數據塊大小由 PAX 配置參數(預設為 60k)在 NAS 上配置。
2016-05-25 17:30:37: NDMP: 3: Thread ndmp533 DartTapeInterface: tape read error: tape read block size (262144 bytes) > 131072 bytes (NDMP.bufsz (128 KB) x 1024)
2016-05-25 17:30:37: NDMP: 3: Thread ndmp533 set parameter NDMP.bufsz (in KB) to 256 or larger to resolve this error
2016-05-25 17:30:49: NDMP: 3: Thread ndmp535 DartTapeInterface: tape read error: tape read block size (262144 bytes) > 131072 bytes (NDMP.bufsz (128 KB) x 1024)
2016-05-25 17:30:49: NDMP: 3: Thread ndmp535 set parameter NDMP.bufsz (in KB) to 256 or larger to resolve this error
2016-05-25 17:30:50: NDMP: 3: Thread ndmp536 DartTapeInterface: tape read error: tape read block size (262144 bytes) > 131072 bytes (NDMP.bufsz (128 KB) x 1024)
2016-05-25 17:30:50: NDMP: 3: Thread ndmp536 set parameter NDMP.bufsz (in KB) to 256 or larger to resolve this error
VNX 記錄在備份時顯示相關訊息:
1466097048: NDMP: 3: Session 236 (thread ndmp236) TAPE_WRITE, size(262144) > ndmpMaxBufSize (131072) Set NDMP.bufsz >= 262144 and reboot
這不被視為備份問題,因為數據流尚未啟動,但它會導致 NMDP 標頭損壞 (請參閱:NDMP:3:工作階段 384 (thread ndmp384) TAPE_WRITE, size(262144) > ndmpMaxBufSize (131072)Set NDMP.bufsz >= 262144 並重新開機),可將 NDMP 裝置緩衝區大小設定為 256k,以解決此問題。但是,此更正不能解決已寫入備份的問題。
上述 NMDP 標頭損壞可能存在而不會導致恢復問題,但在備份重新啟動的情況下,在 NW 媒體資料庫中註冊的檔案編號是錯誤的。這會使磁帶的自動定位失效,因此 NAS 回應說該位置的現有資料無法識別為有效的 NDMP 資料串流 (NDMP 服務錯誤:無法識別格式)
所以這裡有兩種類型的損壞:
- NetWorker 和 VNX 之間的區塊大小設定不一致導致 NDMP 標頭 IO 錯誤,自動定位功能會略過這種類型的損毀,而且它可能存在多年,且不會讓您注意到。只要在每個集區註冊了正確的檔案編號即可。
- 重新啟動備份時,會依序寫入多個 NDMP 標頭和頁腳,並且 NW 伺服器將錯過下一次備份開始的正確檔案編號,這使得無法使用自動定位 NDMP 禁用自動定位調試 (
/nsr/debug/ndmp_auto_pos) 在這種情況下毫無用處,因為磁帶管理介面 (MTIO) 在遇到雙檔標記(由 C1 NDMP 標頭損壞引起)時將停止讀取
Resolution
聯絡 Dell 支援時,請參考此 KB。