Unsolved
This post is more than 5 years old
1 Rookie
•
8 Posts
0
13483
May 21st, 2013 09:00
Recover MCS config from crashed AVE
Hello,
I have a crashed AVE that I need to pull the MCS config from and try to apply it to a new AVE. I can get the crashed one to start the gsan, even with the invalid checkpoints (reason for crash) but I cant get the mcs start because of the gsan errors.
I can still SCP and putty into the crashed system so I was hoping there was a way I could just extract or copy the config files from it and import or copy into our new AVE. The crashed AVE is 6.0.1 and the new one is 6.1, if that matters.
I didn't setup the original AVE so I don't know how the polices, schedules, datasets, etc. were configured. If I could get the mcs to start then I could at least look at both systems and copy the configs piece by piece. If not, then I'm hoping there is a way to get the data out of the crashed one and at least view it so I can make an attempt at configuring the new one.
Can anyone help?
Thanks,
Justin


ionthegeek
2 Intern
•
2K Posts
0
May 21st, 2013 09:00
Was the AVE that crashed replicating elsewhere? If so, support can assist you with replicating and restoring the MCS config to the new AVE. This would likely be faster and more reliable than trying to retrieve this information from the damaged system.
iambuga
1 Rookie
•
8 Posts
0
May 21st, 2013 09:00
ianderson,
Unfortunately, no. This was a stand-alone system.
ionthegeek
2 Intern
•
2K Posts
0
May 21st, 2013 10:00
What happens when you try to start up MCS? Can you paste the output here?
iambuga
1 Rookie
•
8 Posts
0
May 21st, 2013 13:00
Sure. Here is the output:
admin@ohmavamar:~/>: dpnctl status
Identity added: /home/admin/.ssh/dpnid (/home/admin/.ssh/dpnid)
dpnctl: INFO: gsan status: not running
dpnctl: INFO: MCS status: down.
dpnctl: INFO: EMS status: down.
dpnctl: INFO: Backup scheduler status: down.
dpnctl: INFO: dtlt status: down.
dpnctl: INFO: axionfs status: up.
dpnctl: INFO: Maintenance windows scheduler status: suspended.
dpnctl: INFO: Unattended startup status: disabled.
dpnctl: INFO: [see log file "/usr/local/avamar/var/log/dpnctl.log"]
admin@ohmavamar:~/>: dpnctl start gsan
Identity added: /home/admin/.ssh/dpnid (/home/admin/.ssh/dpnid)
dpnctl: INFO: Checking that gsan was shut down cleanly...
- - - - - - - - - - - - - - - - - - - -
Here is the most recent available checkpoint:
Fri May 10 01:19:36 2013 UTC Not Validated
A rollback is recommended: the gsan was not shut down cleanly.
The choices are as follows:
1 roll back to the most recent checkpoint, whether or not validated
2 (not applicable: no validated checkpoints are available)
3 select a specific checkpoint to which to roll back
4 restart, but do not roll back (NOT RECOMMENDED:
gsan reported unclean shutdown)
5 do not restart
q quit/exit
(Entering an empty (blank) line twice quits/exits.)
> 4
dpnctl: INFO: Restarting the gsan (this may take some time)...
dpnctl: INFO: To monitor progress, run in another window: tail -f /tmp/dpnctl-gsan-restart-output-6325
dpnctl: WARNING: 1 warning seen in output of "/usr/bin/yes no | /usr/local/avamar/bin/restart.dpn"
dpnctl: ERROR: 18 errors seen in gsan error logs:
- - - - - - - - - - - - - - - - - - - - BEGIN
(0.0) 2013/05/21-20:25:55.92343 {0.0} [disk6] ERROR: <0001> tstripeheader::validate magic value = 0x0 out of range 0x12345678 .. 0x12345678
(0.0) 2013/05/21-20:25:55.92351 {0.0} [disk6] ERROR: <0001> tstripeheader::validate length value = 0 out of range 4096 .. 4096
(0.0) 2013/05/21-20:25:55.92356 {0.0} [disk6] ERROR: <0001> tstripeheader::validate version value = 0x0 out of range 0x1000000 .. 0x1000000
(0.0) 2013/05/21-20:25:55.92361 {0.0} [disk6] ERROR: <0001> tstripeheader::validate created value = 0x0 out of range 0x3b000000 .. 0xffffffff
(0.0) 2013/05/21-20:25:55.92366 {0.0} [disk6] ERROR: <0001> tstripeheader::validate kind value = 0 out of range 6 .. 6
(0.0) 2013/05/21-20:25:55.92372 {0.0} [disk6] ERROR: <0001> tstripeheader::validate wrong stripe id (none) != 0.0-13F7
(0.0) 2013/05/21-20:25:55.92378 {0.0} [disk6] ERROR: <0001> stripe::init stripe=0.0-13F7 proxy=0.0-13F7 bad stripe header
(0.0) 2013/05/21-20:26:24.57915 {0.0} [disk2] ERROR: <0001> tchunklist::validatelist invalid maxlength=10485760 chunkid=17401 [off=74695949 sz=152 type=recipe4(hints)]
(0.0) 2013/05/21-20:27:16.35453 {0.0} [disk5] ERROR: <0001> tchunklist::validatelist overlapping allocated chunkid1=13646 [off=248292633 sz=7137 type=extatomic] chunkid2=31673 [off=248298227 sz=6233 type=extatomic]
(0.0) 2013/05/21-20:27:16.35467 {0.0} [disk5] ERROR: <0001> tchunklist::validatelist overlapping allocated chunkid1=31673 [off=248298227 sz=6233 type=extatomic] chunkid2=13647 [off=248299770 sz=1392 type=extatomic]
(0.0) 2013/05/21-20:29:56.67130 {0.0} [disk0] ERROR: <0001> tchunklist::validatelist overlapping allocated chunkid1=20834 [off=67199308 sz=24845 type=extatomic] chunkid2=2041 [off=67217224 sz=12097 type=atomic]
(0.0) 2013/05/21-20:29:56.67139 {0.0} [disk0] ERROR: <0001> tchunklist::validatelist overlapping allocated chunkid1=2041 [off=67217224 sz=12097 type=atomic] chunkid2=20835 [off=67224153 sz=10711 type=extatomic]
(0.0) 2013/05/21-20:30:32.42376 {0.0} [disk5] ERROR: <0001> tchunklist::validatelist overlapping allocated chunkid1=29601 [off=361816356 sz=6458 type=extatomic] chunkid2=24057 [off=361822028 sz=10955 type=extatomic]
(0.0) 2013/05/21-20:30:32.42384 {0.0} [disk5] ERROR: <0001> tchunklist::validatelist overlapping allocated chunkid1=24057 [off=361822028 sz=10955 type=extatomic] chunkid2=29602 [off=361822814 sz=15723 type=extatomic]
(0.0) 2013/05/21-20:30:36.44178 {0.0} [disk11] ERROR: <0001> tchunklist::validatelist overlapping allocated chunkid1=14999 [off=91077795 sz=20523 type=extatomic] chunkid2=7097 [off=91079860 sz=19497 type=extatomic]
(0.0) 2013/05/21-20:30:36.44187 {0.0} [disk11] ERROR: <0001> tchunklist::validatelist overlapping allocated chunkid1=7097 [off=91079860 sz=19497 type=extatomic] chunkid2=15000 [off=91098318 sz=36559 type=extatomic]
(0.0) 2013/05/21-20:31:05.84126 {0.0} [disk2] ERROR: <0001> tchunklist::validatelist overlapping allocated chunkid1=5532 [off=224898136 sz=14356 type=extatomic] chunkid2=8185 [off=224902638 sz=37755 type=extatomic]
(0.0) 2013/05/21-20:31:05.84135 {0.0} [disk2] ERROR: <0001> tchunklist::validatelist overlapping allocated chunkid1=8185 [off=224902638 sz=37755 type=extatomic] chunkid2=5533 [off=224912492 sz=14271 type=extatomic]
- - - - - - - - - - - - - - - - - - - - END
dpnctl: INFO: Restarting gsan succeeded.
dpnctl: INFO: gsan started.
dpnctl: INFO: [see log file "/usr/local/avamar/var/log/dpnctl.log"]
admin@ohmavamar:~/>: dpnctl start mcs
Identity added: /home/admin/.ssh/dpnid (/home/admin/.ssh/dpnid)
dpnctl: INFO: Starting MCS...
dpnctl: INFO: To monitor progress, run in another window: tail -f /tmp/dpnctl-mcs-start-output-8845
dpnctl: ERROR: error return from "[ -r /etc/profile ] && . /etc/profile ; /usr/local/avamar/bin/mcserver.sh --start" - exit status 1
dpnctl: ERROR: 1 error seen in output of "[ -r /etc/profile ] && . /etc/profile ; /usr/local/avamar/bin/mcserver.sh --start"
dpnctl: INFO: [see log file "/usr/local/avamar/var/log/dpnctl.log"]
admin@ohmavamar:~/>:
ionthegeek
2 Intern
•
2K Posts
1
May 21st, 2013 13:00
It looks like GSAN was rolled back on its own but the MCS was never restored after. It's vitally important for the GSAN user accounting system and the MCS database to be in sync which is why the system disallows MCS restart after rollback.
You can restore the MCS using dpnctl start mcs --force_mcs_restore if you just want to get the MCS back up and take a look at the settings.
In any case, if the AVE is trashed anyway, a replicate of the most recent MCS flushes to the new system and an MCS restore will probably get you back on your feet the fastest. Support can assist with this process so I would recommend opening a service request.
iambuga
1 Rookie
•
8 Posts
0
May 21st, 2013 13:00
And the error in the dpnctl log is:
ERROR: gsan rollbacktime: 1368741588 does not match stored rollbacktime: 1368142479
ionthegeek
2 Intern
•
2K Posts
0
May 21st, 2013 14:00
The MCS configuration files may be removed if a flush restore fails and this will lead to future attempts failing. It is possible to manually extract the missing files from a recent flush but this is somewhat error prone so I would strongly recommend working with support for this issue.
iambuga
1 Rookie
•
8 Posts
0
May 21st, 2013 14:00
I tried doing the --force_mcs_restore but it still errors out and says that the restore did not succeed so its not restarting MCS.
I'm going to look through the file system and see if I can find any of the MCS backups and maybe SCP them to the other server. I cant use any of the migration tools because I cant get the MCS server up.
ionthegeek
2 Intern
•
2K Posts
0
May 23rd, 2013 14:00
No there is not.
Edit: Sneaky edit while I was composing my reply! I saw that
Copying checkpoints around isn't supported. The ability to back up checkpoints from single nodes to DD is planned for the upcoming release (standard disclaimer about features and release dates being subject to change, etc.).
niko.virta
51 Posts
0
May 23rd, 2013 14:00
Btw. Is there documented way to backup MCS on single node instance which is not replicated? ie. doing postgres dump and transfer it and maybe some files via scp/sftp to another location? Getting all the data for schedules, datasets, domains etc. would help enormously if starting from scratch.
I guess you could transfer validated checkpoint also to somewhere else? 7.0 hopefully brings GSAN backup to DD which would be nice feature.
pannujat
3 Posts
0
May 11th, 2015 22:00
Worked!