I am working on restoring large amounts of data from snapshotIQ snapshots. Many of the directories are not 'huge' in the sense of size, but are massive in terms of number of files; ranging anywhere from a few thousand files, to the extreme of 500+ million files. Because of this, these restores take a VERY long time.
Unfortunately, for most of these the snapshot was taken at the parent level (something I am going to re-evaluate, lesson learned), so I cannot use SnapRevert because I dont want to revert every sub-directory under the snapshot.
So I am left with copying the directories manually out of the snapshots. (/ifs/data/.snapshot/....)
looking at the Isilon's cp man page, there are two OneFS unique switches, -c (clones) and -B (use large buffers). I did some testing with -c with a test directory and it appears that I cannot use -c and -R (recursive) together, so that doesn't appear to be an option for me.
I am curious about -B, it says "Use larger buffers when copying data to regular files to reduce system calls". I'd like to know the full impact if I use -B; will it speed up my copy jobs with millions of 'smallish' files? What's the impact on the cluster?
alternatively, rsync is another option and is obviously more robust than cp, but I am concerned about time and not sure if rsync would be faster.
thanks in advance for any information/recommendations
You can use SyncIQ to replicate out of a snapshot, but you can only do it on the command line.
If you do not have a SyncIQ licenses contact your Local Systems Engineer and they can provide you a demo license
Example: If you wanted to replicate /ifs/data/.snapshot/snapshotname/dir-with-millions-of-files
On the cli you would use a job like this "isi sync job start --policy-name=name-of-synciq-policy --source-snapshot=source-snapshot-name"Note:
Note: you do have to have a syncIQ policy setup for the directory in question, you simply are changing the root of the replication for the job to the snapshot. IE: in the above example you would have a sync job for /ifs/data/dir-millions-of-files
Your local SE can assist you with this, this is a very efficient way to recover large amounts of data from within snapshots.
I actually tried syncIQ, we are licensed and use it quite a bit to replicate to our other cluster. However, when I setup a policy to replicate from the snapshot directory back to the original location on the local cluster, it failed when i started the job stating that it failed to snapshot the source, which makes sense for how synciq works.
So I guess that's where the command line argument you mentioned comes into play. So to clarify your recommendation, when you say i have to have a syncIQ policy setup for the directory in question. Do you mean to setup a policy like what i described above (source=snapshot, dest=original location), just only run it from cli with that snapshot argument in it?
otherwise, if i had a synciq job setup like my normal stuff (source cluster --> target cluster), it seems that if i ran that policy with the snapshot source argument, it would simply replicate the snapshots to the remote target cluster.
Actually you need a normal syncIQ policy (Just like you would configure without a snapshot)
Then use the CLI flat to define the snapshot to replicate from.
This is essentially a change root for the single run of the SyncIQ Job.
You cannot configure the snapshot into the policy, or set source=snapshot.
What you setup is to replicate an entire snapshot, and that is not what you want you want to replicate a subset of a snapshot so the CLI setup is the only way.
thank you for the info. Though, I guess I need further clarification regarding the 'regular' policy that needs to be setup.
Since this is not normally a directory that I would have replicating locally, what am I defining as the 'source' in this regular policy?
for example, let's say /ifs/data/directory/directory_1 is the directory I am trying to restore to (the original location). I would obviously set this up as the destination, but other than the snapshot, there's no 'source' that contains this data. Would I just use a miscellaneous empty directory as source since the snapshot will be swapped in as source via --source-snapshot on CLI?
i.e.
source= /ifs/data/empty_dir
dest = /ifs/data/directory/directory_1
but then run it with "isi sync job start --policy-name=temp_policy --source-snapshot=/ifs/.snapshot/snapname/data/directory/directory_1"?
thanks, i have successfully tested this now. that was my missing piece, that I could not go back to the original location.
So I have a policy with the source as the original location and a new location as the destination. run the synciq policy with the snapshot specified with --source-snapshot, and I got exactly the data I needed.
I can then rename the original directory to something else with mv, and then rename the restore location to the original name.
Peter_Sero
6 Operator
•
1169 Posts
5297
0
Posted September 6th, 2017 12:00
If you have removed the directory on which the original snapshot was taken,
it appears to me that you will be out of luck with using SyncIQ for restoring.
If that directory still exists, and you want to restore only some subdirectory,
this can be specified with
isi sync policies {create or modify} ... --source-include-directories /path/to/subdirectory
For local SyncIQ jobs, the target path must be outside the source path,
so one cannot restore a lost directory "in-place".
hth
-- Peter