Norconex Crawler v4 Beta 2: A Security Fix, and Content Ready for AI
Beta 2 is here. It is mostly what a second beta should be: fixes earned by the first few weeks of real use, not new surface area. Two things are worth your attention regardless โ one you should act on, one you can put to work.
If you're running beta.1, upgradeโ
The crawler's cluster administrative server starts with every crawl, not only clustered ones, and in beta.1 it bound every network interface by default. Its stop endpoint is guarded by a crawler ID โ a configuration value, not a secret โ so anyone able to reach the port could halt a crawl.
It now binds to loopback only by default. If you run a clustered deployment
that needs the admin server reachable from another host, set
cluster.adminBindAddress
explicitly (any, or a specific address) โ the crawler logs a warning
whenever it binds beyond loopback, so that choice can't pass unnoticed.
Prepare content for AI, during the crawlโ
Two new Importer handlers:
- TextChunkSplitter cuts long documents into pieces sized for embedding models, breaking at paragraph or sentence boundaries rather than mid-word.
- TextEmbeddingTransformer calls any OpenAI-compatible embeddings endpoint โ OpenAI itself, a local Ollama, a LiteLLM proxy, vLLM, and anything else speaking that same request shape โ and stores the resulting vector as a document field.
Chain them and each chunk becomes its own small document, embedded on the
way through, ready for a vector-capable index. Elasticsearch, OpenSearch,
and Solr can already hold that vector: list the field in the
Elasticsearch or
OpenSearch committer's jsonFieldsPattern so it is sent as numbers instead
of a quoted string, or let Solr's DenseVectorField take it as-is.
Results are cached by a hash of the text, the model, and the endpoint, so the same content is never billed twice โ whether it recurs on a later crawl, or as boilerplate repeated across many pages in the same one. See the AI section of the Features page for where this fits next to everything else the crawler already does for retrieval and graph pipelines.
Also in this releaseโ
- Example configurations, fixed. The bundled web crawler
minimumexample was still V3-shaped and failedconfigcheckoutright โ the first thing a new user runs, broken. Bothminimumandcomplexare rewritten for V4 and now point at dedicated pages built for the purpose, instead of crawling the corporate site with onlymaxDepthas a brake. The File System Crawler had no examples at all before this release; it now ships runnable ones alongside a refreshed HOWTO, and defaults toLocalFetcherso a minimal configuration works out of the box. - Orphan handling: references that were never actually committed are no longer deleted as orphans, and orphans are no longer deleted at all when a crawl ends early โ both could previously drop documents from your index that were never confirmed as gone.
- A null pointer fetching a site's
robots.txt/sitemap no longer crashes a crawl. - Event-count reporting now accounts for every event type across a whole session, and resets correctly when a new one starts.
The full changelog has the complete list.
Get itโ
No configuration changes are required to move from beta.1. Grab the update from Download, or pull the updated Docker image โ same as before, Java 21 is the only prerequisite.
Still a beta: expect rough edges, and the occasional breaking change before 4.0.0 is final. Bugs, questions, and ideas are all welcome on GitHub Discussions.