We recently upgraded a cluster of Isilon 108NL nodes (OneFS 7.0.1.2) by adding six NL400 nodes to it, which then formed their own node pool. Since we want to retire the old 108NL nodes as part of the upgrade, we ran a SmartPools job whose filter simply was
[File name matches "*"]
and whose operation was to
[Save data to "NL400 pool"]
and it successfully moved all the data, hundreds of terabytes, to the new nodes.
However, SmartPools jobs continue to run against data that is, for some reason, still collecting in the old node pool. Why? Could it be, that, once SmartPools walks a particular part of the file system tree, and it records which files need to be moved, and then goes to another part of the tree to take further inventory, any new files written get skipped, until the job runs again?
A recent discussion suggests that the use of a folder designation in the filter, rather than "*", may work better. Would it? For that matter, what is the best practice for directing any newly written, or "in flight" data, to the new NL400 nodes only? Any advice appreciated.
Could it be, that, once SmartPools walks a particular part of the file system tree, and it records which files need to be moved, and then goes to another part of the tree to take further inventory, any new files written get skipped, until the job runs again?
No, it doesn't work that way.
A recent discussion suggests that the use of a folder designation in the filter, rather than "*", may work better. Would it?
No, this should work just as well.
For that matter, what is the best practice for directing any newly written, or "in flight" data, to the new NL400 nodes only?
What you did sounds pretty appropriate to me.
I'm not sure exactly what you mean by "still collecting in the old node pool". Are you noticing specific files, or merely seeing nonzero utilization? Once the file pool policy was changed to set everything to the NL400 and the SmartPools job finished, all file data should be on the NL400 pool. There are a couple of potential exceptions why that might not be the case including files or directories being manually managed or the SmartPools job not having completed yet, or system data such as the LIN tree being stored on the old nodes. If you're aware of an example of a file that should be on NL400 according to the file pool policy but isn't, you can examine its attributes to help determine what's going on.
Do you know whether there are any files or directories on the cluster which are (intentionally) manually managed? If not, the easiest solution is probably to tell SmartPools to override that and move manually managed files as well.
Using SmartPools for migration will make the eventual removal of the 108NLs faster, but once they're down to very low utilization, you can SmartFail them. It's not like they have to be to absolute 0 before that.
Your rule will move the files, however if you do not change your default policy to point all new writes to the NL400 nodes, you will likely find that your old nodes continue to get targeted for writes. You likely have the default policy being to write to the ANY node pool. You should change it to write to the NL400 node pool and this should solve your problem.
Your file name match filter will only run when SmartPools itself runs. If you instead targeted a directory path to live on the NL400 pool, then that would take effect on all new writes as well.
Just to be clear, the default file pool policy isn't immediately applied. It requires a SmartPools (or SetProtectPlus) job to run just like user-defined file pool policies. If a rule matching "*" is first in the list, it will apply to all files and the default file pool policy shouldn't matter (except potentially for attributes not specified by the "*" file pool policy).
Thanks for all the helpful insight so far. Some additional information has come to light, as we try to determine why approximately 2TB of "data" is reported as still residing in the old pool. It does not look like this data was "in flight" after all. A look at the SmartPools page shows that 1.98 TB, or .34%, is stubbornly staying behind, in the pool 'iq_108NL'. File System Explorer does not seem to indicate that this data is being manually managed. And finally, a log file of the last SmartPools job to run is appended. Could the 1.8TB be metadata? Is this normal? Some questions:
1) FS Explorer doesn't list hidden folders. How to we find out about them?
2) If there *were* any manually managed data, how would we find that out what the files were, and their attributes, beyond drilling down, and down, (and down!) using FS Explorer. Is there some kind of 'show all manually managed data' functionality?
3) If files (data) were open by an application writing to them, might those files refuse to be moved by a SmartPools job as well?
Screen shots are appended. We can of course open a case with tech support, but wonder if there is not some logical explanation we can come to without having to do that.
Thanks,
Don
here is the log from the last SmartPools run:
>
> primary-10# isi job history --job 20350 --verbose
1.8TB is less than 0.5% of the total data. That is quite likely metadata.
FS Explorer doesn't list hidden folders. How to we find out about them?
isi get should work.
If there *were* any manually managed data, how would we find that out what the files were, and their attributes, beyond drilling down, and down, (and down!) using FS Explorer. Is there some kind of 'show all manually managed data' functionality?
No, this is why manually managed data is such a pain and should be avoided. It undermines the simplicity of management.
If files (data) were open by an application writing to them, might those files refuse to be moved by a SmartPools job as well?
The job would wait, I believe.
Another possibility is that some of that data is leaked blocks. It's such a small percentage, it doesn't seem disconcerting to me. Is there any reason not to just go ahead with the smartfail of the 108NLs?
jbauman
3 Posts
1544
0
Posted January 6th, 2015 10:00
No, it doesn't work that way.
No, this should work just as well.
What you did sounds pretty appropriate to me.
I'm not sure exactly what you mean by "still collecting in the old node pool". Are you noticing specific files, or merely seeing nonzero utilization? Once the file pool policy was changed to set everything to the NL400 and the SmartPools job finished, all file data should be on the NL400 pool. There are a couple of potential exceptions why that might not be the case including files or directories being manually managed or the SmartPools job not having completed yet, or system data such as the LIN tree being stored on the old nodes. If you're aware of an example of a file that should be on NL400 according to the file pool policy but isn't, you can examine its attributes to help determine what's going on.
Do you know whether there are any files or directories on the cluster which are (intentionally) manually managed? If not, the easiest solution is probably to tell SmartPools to override that and move manually managed files as well.
Using SmartPools for migration will make the eventual removal of the 108NLs faster, but once they're down to very low utilization, you can SmartFail them. It's not like they have to be to absolute 0 before that.