Skip to main content

File System Crawler Quick Start

This guide gets you from zero to a running file system crawl in under 5 minutes.

Just want to see it run?

The distribution ZIP ships ready-to-run examples that crawl bundled sample files offline. See Run the bundled examples to try one before writing your own config.

Built-in baseline, not a hard limit

The source protocols and examples in this guide represent built-in support in Norconex Crawler v4. They are practical defaults, not a fixed ceiling. Teams can extend crawler behavior with custom connectors, parser logic, and pipeline components for customer-specific requirements.

Windows users

Replace crawl-fs.sh with crawl-fs.bat and ./ with .\ in all commands below.

Step 1 โ€” Create a config fileโ€‹

Create my-fs-crawl.yaml with the following minimal configuration:

id: my-first-crawl
numThreads: 5
startReferences:
- /path/to/your/documents
committers:
- class: LogCommitter
ignoreContent: true
logLevel: INFO

This crawl will process every file under the given path and log each document (LogCommitter is built-in and ideal for testing before connecting a real backend).

Step 2 โ€” Filter by file type (optional)โ€‹

To crawl only specific file extensions:

id: my-first-crawl
startReferences:
- /path/to/your/documents
referenceFilters:
- class: ExtensionReferenceFilter
onMatch: INCLUDE
extensions:
- pdf
- docx
- xlsx
- pptx
- txt
- html

Step 3 โ€” Start the crawlโ€‹

./crawl-fs.sh start -config=my-fs-crawl.yaml

You'll see log output as files are fetched, filtered, and committed.

Docker alternativeโ€‹

If you prefer running with Docker, mount your config and logs, then run:

docker run --rm \
-v "${PWD}:/opt/norconex/crawler/configs" \
-v "${PWD}/logs:/opt/norconex/crawler/logs" \
-e COLLECTOR_CONFIG_FILE=my-fs-crawl.yaml \
norconex/crawler-fs:latest
During the v4 beta

No latest tag is published until 4.0.0 ships. Use the current pre-release tag instead (for example norconex/crawler-fs:4.0.0-beta-1); the download page always shows it.

For Docker Compose examples and release-tag guidance, see Docker.

Step 4 โ€” Stop and resumeโ€‹

Stop the crawl at any time:

./crawl-fs.sh stop -config=my-fs-crawl.yaml

Norconex saves its state automatically. Resume exactly where you left off by running the same start command again:

./crawl-fs.sh start -config=my-fs-crawl.yaml

To clear the crawler state before your next run, you can issue this command:

./crawl-fs.sh clean -config=my-fs-crawl.yaml

Alternatively, you can combine "clean" with the start command:

./crawl-fs.sh start -clean -config=my-fs-crawl.yaml

Step 5 โ€” Remote file systemsโ€‹

The file system crawler supports remote protocols out of the box. A few examples:

If you are unsure which scheme or start reference format to use for a given fetcher, see FS Fetchers Quickstart.

SFTP:

id: my-first-crawl
startReferences:
- sftp://fileserver.example.com/data/documents
credentials:
username: user
password: myPassword

WebDAV (Nextcloud, older versions of SharePoint):

id: my-first-crawl
startReferences:
- https://sharepoint.example.com/sites/mysite/Shared Documents/
credentials:
username: user
password: myPassword

Apache HDFS:

id: my-first-crawl
startReferences:
- webhdfs://namenode:9870/user/data/corpus

For Box, Google Drive, Egnyte, and M365-specific examples, use the dedicated fetcher reference pages so configuration examples stay canonical in one place:

CMIS is a standard, so one fetcher covers Alfresco, Nuxeo, OpenText, Documentum, IBM FileNet and any other compliant repository. See Content Sources for the full list.

Step 6 โ€” Send to a real targetโ€‹

Replace the LogCommitter with your actual destination. See the Integrations page for all available committers and their configuration.

When using ZIP distributions, external committers are downloaded separately as nx-committer-<name>-<version>.zip and their lib/*.jar files must be copied into the crawler lib/ directory. Built-in committers such as LogCommitter do not require this extra step.

CLI referenceโ€‹

CommandDescription
start -config=<file>Start or resume a crawl
stop -config=<file>Gracefully stop a running crawl
clean -config=<file>Delete crawl state (forces a full recrawl)

Next stepsโ€‹