Installation
Requirements
- Java 21 or later (JRE is sufficient for CLI use)
Verify your Java version:
java -version
Standard vs Full distributions
Each crawler ships in two variants:
- Standard — covers most use cases and has a smaller footprint.
- Full — includes everything in Standard plus optional extras: OCR, NLP language detection, extra parsers, scripting, and more.
Start with Standard unless you know you need those extras.
Option 1 — Download the ZIP
-
Go to the Download page and grab the latest release ZIP for your chosen crawler.
-
Extract it (the folder name matches the ZIP name):
# Standard distributionunzip nx-crawler-web-4.0.0-standard.zip# Full distributionunzip nx-crawler-web-4.0.0-full.zip -
The extracted folder looks like this:
nx-crawler-web-4.0.0-standard/crawl-web.sh ← Linux/macOS launch scriptcrawl-web.bat ← Windows launch scriptlog4j2.xml ← logging configurationexamples/ ← ready-to-run sample configurationslib/ ← all JAR dependenciesscripts/ ← utility scriptsthird-party/ ← third-party noticesREADME.txtLICENSE.txt -
On Linux/macOS, make the launch script executable:
chmod +x crawl-web.sh -
Verify the installation:
# Linux/macOS./crawl-web.sh --help# Windows.\crawl-web.bat --help
All examples above use the Web Crawler (nx-crawler-web, crawl-web.sh).
The File System Crawler is installed identically — substitute web with fs in every filename
and command (e.g. nx-crawler-fs-4.0.0-standard.zip, crawl-fs.sh).
Install external committers when using ZIP distributions
ZIP distributions include built-in committer support from nx-committer-core
(for example, LogCommitter).
If you want an external committer such as Elasticsearch, Solr, SQL, Kafka, Neo4j,
Amazon CloudSearch, or Azure Cognitive Search, download the matching
nx-committer-<name>-<version>.zip package separately and copy its JAR files into
the crawler lib/ directory.
Example (Linux/macOS):
# From the crawler installation directory
unzip -j nx-committer-elasticsearch-4.0.0.zip "*/lib/*.jar" -d lib/
Example (Windows PowerShell):
# From the crawler installation directory
Expand-Archive .\nx-committer-elasticsearch-4.0.0.zip -DestinationPath .\tmp-committer
Copy-Item .\tmp-committer\*\lib\*.jar -Destination .\lib\ -Force
Remove-Item .\tmp-committer -Recurse -Force
After copying the JARs, start (or restart) the crawler normally.
Option 2 — Use Docker images
Norconex publishes official images to both Docker registries:
- Docker Hub:
norconex/crawler-web,norconex/crawler-fs,norconex/crawler-web-playwright - GitHub Container Registry:
ghcr.io/norconex/crawler-web,ghcr.io/norconex/crawler-fs,ghcr.io/norconex/crawler-web-playwright
Registry pages:
Example pulls:
docker pull norconex/crawler-web:latest
docker pull ghcr.io/norconex/crawler-fs:latest
Version 4.0.0 has not been released yet, so there is no latest tag published
yet. Until it ships, use the current pre-release tag instead (for example
norconex/crawler-web:4.0.0-beta-1) — the download page always
shows the current one.
Official Docker images already bundle the external committer JARs, so no separate committer installation is required in containers.
For Docker Compose examples, tag strategy (latest vs edge), and runtime tips,
see Docker.
Option 3 — Maven dependency
For embedding the crawler in a Java application, see Java Integration.
Configuration rules and defaults
Before authoring larger configs, read Configuration Semantics for shared behavior such as defaults, null vs empty values, variable resolution, and reusable fragments.
Option 4 — Build from source
git clone https://github.com/Norconex/crawler.git
cd crawlers
# Web crawler
mvn clean package -pl crawler/web -Dmaven.test.skip=true
# File system crawler
mvn clean package -pl crawler/fs -Dmaven.test.skip=true
The distribution ZIP files will be in crawler/web/target/ or crawler/fs/target/.
Next: Docker, Web Quick Start, or File System Quick Start