WebSource Harvester is an educational web-source harvester that crawls a site (BFS, depth-controlled), downloads browser-visible assets (HTML, CSS, JS, images, fonts, PDFs), and rewrites paths so pages work offline, including nested routes. It enforces same-origin limits and is designed for learning, offline analysis, and safe portfolio demos.
javascript css python html pdf fonts crawler offline images web-scraping robots educational bfs srcset student-project same-origin github-actions authorized-testing site-mirroring path-rewriting
-
Updated
Mar 2, 2026 - Python