Summarize web articles: extract clean content, title, image, language, and generate text summaries
-
Updated
Jul 19, 2026 - Python
Summarize web articles: extract clean content, title, image, language, and generate text summaries
open-news is a Python library for news discovery and extraction. It fetches live news via DuckDuckGo, searches Google News, discovers RSS feeds, crawls websites, and extracts article text. Features filtering, deduplication, ranking, and optional JavaScript rendering. Includes Python API, CLI, and terminal UI.
This repository is part of an NLP course for humanities and cultural studies. This course uses historical newspapers as a source and applies NLP methods to them. NLP tasks: Tokenization, Lemmatization, TF-IDF, Part-of-speech tagging, semantic search with transformers, article extraction and OCR post-correction with LLMs, NER and text classification
GNewsScraper is a TypeScript package that scrapes article data from Google News based on a keyword or phrase. It returns the results as an array of JSON objects, making it convenient to access and use the scraped information
强大易用的文章转 Markdown 工具,一键采集微信公众号、掘金、CSDN 等文章,自动下载图片、清理代码块、支持批量转换和 Agent Skill
A simple HTML-to-Markdown converter with article extraction, selector filtering, and batch conversion.
OpenClaw Skill:读取微信公众号文章、识别公众号并拉取文章列表
📋 WebMD is a Chrome extension that transforms web pages into Markdown documents with surgical precision.
A configurable pipeline for extracting and filtering articles from large corpora, tailored for the Delpher Kranten corpus, with support for features like keyword filtering and tf-idf-based relevance scoring.
Mantis grabs exactly what you'd see on the page and turns it into structured JSON or clean Markdown. I built it for read-later and bookmarking tools, and for AI agents that need input without all the token bloat. It's Readability-style, has zero dependencies, and lives in a single file.
A pipe-based news article scraping and metadata extraction library for Python
Zendesk articles extraction toolkit
HTML main-content extraction for Rust — ports of Mozilla Readability, Trafilatura, and htmldate.
Chrome/Edge extension that estimates article word count, reading time, and lets you double-click any word to update the toolbar badge with your reading progress.
Chrome extension that yoinks webpages into clean markdown. Supports article extraction, full-page capture, YouTube transcripts, and visual element picking.
AYLIEN — independent third-party profile of a public API surface, by API Evangelist. AYLIEN is a news intelligence and text analysis platform providing REST APIs for article extraction, sentiment analysis, entity recognition, summarization, concept detection, and NLP-enriched news aggregation from over 80,000 public and licensed sources delivering
India Times news extraction tool
Production web scraper with Playwright, bot-detection plugins, fingerprint rotation, and CAPTCHA solving. CLI + FastAPI.
Pure-Python article extraction library and HTTP API - Extract clean content from web pages as Markdown or HTML
To associate your repository with the article-extraction topic, visit your repo's landing page and select "manage topics."