A scalable, mature and versatile web crawler based on Apache Storm
-
Updated
Sep 3, 2026 - Java
A scalable, mature and versatile web crawler based on Apache Storm
Resources for running StormCrawler with Docker services
Process web archives (WARC format) with StormCrawler and index content into OpenSearch
Ansible playbook for deploying a Storm cluster
StormCrawler topology to evaluate the performance of different backends and configurations
Crawling pipeline to estimate the adoption of the carbon.txt protocol
Consumer-side conformance probe for StormCrawler -> TikaV2.ParseBytes -> typed Document
To associate your repository with the stormcrawler topic, visit your repo's landing page and select "manage topics."