Skip to main content

Norconex Crawler v4 Beta 2: A Security Fix, and Content Ready for AI

ยท 4 min read
Norconex Team
Norconex Team
Norconex Inc.

Beta 2 is here. It is mostly what a second beta should be: fixes earned by the first few weeks of real use, not new surface area. Two things are worth your attention regardless โ€” one you should act on, one you can put to work.

If you're running beta.1, upgradeโ€‹

The crawler's cluster administrative server starts with every crawl, not only clustered ones, and in beta.1 it bound every network interface by default. Its stop endpoint is guarded by a crawler ID โ€” a configuration value, not a secret โ€” so anyone able to reach the port could halt a crawl.

It now binds to loopback only by default. If you run a clustered deployment that needs the admin server reachable from another host, set cluster.adminBindAddress explicitly (any, or a specific address) โ€” the crawler logs a warning whenever it binds beyond loopback, so that choice can't pass unnoticed.

Prepare content for AI, during the crawlโ€‹

Two new Importer handlers:

  • TextChunkSplitter cuts long documents into pieces sized for embedding models, breaking at paragraph or sentence boundaries rather than mid-word.
  • TextEmbeddingTransformer calls any OpenAI-compatible embeddings endpoint โ€” OpenAI itself, a local Ollama, a LiteLLM proxy, vLLM, and anything else speaking that same request shape โ€” and stores the resulting vector as a document field.

Chain them and each chunk becomes its own small document, embedded on the way through, ready for a vector-capable index. Elasticsearch, OpenSearch, and Solr can already hold that vector: list the field in the Elasticsearch or OpenSearch committer's jsonFieldsPattern so it is sent as numbers instead of a quoted string, or let Solr's DenseVectorField take it as-is.

Results are cached by a hash of the text, the model, and the endpoint, so the same content is never billed twice โ€” whether it recurs on a later crawl, or as boilerplate repeated across many pages in the same one. See the AI section of the Features page for where this fits next to everything else the crawler already does for retrieval and graph pipelines.

Also in this releaseโ€‹

  • Example configurations, fixed. The bundled web crawler minimum example was still V3-shaped and failed configcheck outright โ€” the first thing a new user runs, broken. Both minimum and complex are rewritten for V4 and now point at dedicated pages built for the purpose, instead of crawling the corporate site with only maxDepth as a brake. The File System Crawler had no examples at all before this release; it now ships runnable ones alongside a refreshed HOWTO, and defaults to LocalFetcher so a minimal configuration works out of the box.
  • Orphan handling: references that were never actually committed are no longer deleted as orphans, and orphans are no longer deleted at all when a crawl ends early โ€” both could previously drop documents from your index that were never confirmed as gone.
  • A null pointer fetching a site's robots.txt/sitemap no longer crashes a crawl.
  • Event-count reporting now accounts for every event type across a whole session, and resets correctly when a new one starts.

The full changelog has the complete list.

Get itโ€‹

No configuration changes are required to move from beta.1. Grab the update from Download, or pull the updated Docker image โ€” same as before, Java 21 is the only prerequisite.

Still a beta: expect rough edges, and the occasional breaking change before 4.0.0 is final. Bugs, questions, and ideas are all welcome on GitHub Discussions.