chrstahl89 chrstahl89 - 10 months ago 122
Apache Configuration Question

Dump all segments from nutch

I am simply trying to dump my segments from a crawl using readseg. If I only have one folder the command

bin/nutch readseg -dump crawl/segments/* dumpFolder

works, but if I have multiple segment folders it fails. Any ideas?

Answer Source

You should give the path of segment till the segments dir(the one with the timestamp). If you want to read all the segments in the segments/ dir, you could have a wrapper class where you can list contents in the segments dir and call readseg from there.