File System Crawler Quick Start
This guide gets you from zero to a running file system crawl in under 5 minutes.
The distribution ZIP ships ready-to-run examples that crawl bundled sample files offline. See Run the bundled examples to try one before writing your own config.
The source protocols and examples in this guide represent built-in support in Norconex Crawler v4. They are practical defaults, not a fixed ceiling. Teams can extend crawler behavior with custom connectors, parser logic, and pipeline components for customer-specific requirements.
Replace crawl-fs.sh with crawl-fs.bat and ./ with .\ in all commands below.
Step 1 โ Create a config fileโ
Create my-fs-crawl.yaml with the following minimal configuration:
id: my-first-crawl
numThreads: 5
startReferences:
- /path/to/your/documents
committers:
- class: LogCommitter
ignoreContent: true
logLevel: INFO
This crawl will process every file under the given path and log each document
(LogCommitter is built-in and ideal for testing before connecting a real backend).
Step 2 โ Filter by file type (optional)โ
To crawl only specific file extensions:
id: my-first-crawl
startReferences:
- /path/to/your/documents
referenceFilters:
- class: ExtensionReferenceFilter
onMatch: INCLUDE
extensions:
- pdf
- docx
- xlsx
- pptx
- txt
- html
Step 3 โ Start the crawlโ
./crawl-fs.sh start -config=my-fs-crawl.yaml
You'll see log output as files are fetched, filtered, and committed.
Docker alternativeโ
If you prefer running with Docker, mount your config and logs, then run:
docker run --rm \
-v "${PWD}:/opt/norconex/crawler/configs" \
-v "${PWD}/logs:/opt/norconex/crawler/logs" \
-e COLLECTOR_CONFIG_FILE=my-fs-crawl.yaml \
norconex/crawler-fs:latest
No latest tag is published until 4.0.0 ships. Use the current pre-release tag
instead (for example norconex/crawler-fs:4.0.0-beta-1); the
download page always shows it.
For Docker Compose examples and release-tag guidance, see Docker.
Step 4 โ Stop and resumeโ
Stop the crawl at any time:
./crawl-fs.sh stop -config=my-fs-crawl.yaml
Norconex saves its state automatically. Resume exactly where you left off by running the same start command again:
./crawl-fs.sh start -config=my-fs-crawl.yaml
To clear the crawler state before your next run, you can issue this command:
./crawl-fs.sh clean -config=my-fs-crawl.yaml
Alternatively, you can combine "clean" with the start command:
./crawl-fs.sh start -clean -config=my-fs-crawl.yaml
Step 5 โ Remote file systemsโ
The file system crawler supports remote protocols out of the box. A few examples:
If you are unsure which scheme or start reference format to use for a given fetcher, see FS Fetchers Quickstart.
SFTP:
id: my-first-crawl
startReferences:
- sftp://fileserver.example.com/data/documents
credentials:
username: user
password: myPassword
WebDAV (Nextcloud, older versions of SharePoint):
id: my-first-crawl
startReferences:
- https://sharepoint.example.com/sites/mysite/Shared Documents/
credentials:
username: user
password: myPassword
Apache HDFS:
id: my-first-crawl
startReferences:
- webhdfs://namenode:9870/user/data/corpus
For Box, Google Drive, Egnyte, and M365-specific examples, use the dedicated fetcher reference pages so configuration examples stay canonical in one place:
CMIS is a standard, so one fetcher covers Alfresco, Nuxeo, OpenText, Documentum, IBM FileNet and any other compliant repository. See Content Sources for the full list.
Step 6 โ Send to a real targetโ
Replace the LogCommitter with your actual destination. See the
Integrations page for all available committers and their configuration.
When using ZIP distributions, external committers are downloaded separately as
nx-committer-<name>-<version>.zip and their lib/*.jar files must be copied into
the crawler lib/ directory. Built-in committers such as LogCommitter do not
require this extra step.
CLI referenceโ
| Command | Description |
|---|---|
start -config=<file> | Start or resume a crawl |
stop -config=<file> | Gracefully stop a running crawl |
clean -config=<file> | Delete crawl state (forces a full recrawl) |
Next stepsโ
- Use the Visual Configurator to build your config visually
- Read Concepts: Crawl Pipeline to understand how documents are processed
- Read Concepts: Sessions and Runs to understand resume, deduplication, recrawl policy, and external scheduling