Production-grade genpark-bpe-byte-pair-encoding-tokenizer-skill skill for AI agents
-
Updated
Sep 10, 2026 - Python
Production-grade genpark-bpe-byte-pair-encoding-tokenizer-skill skill for AI agents
Production-grade genpark-bpe-byte-pair-encoding-tokenizer-skill skill for AI agents
Automated data ingestion pipeline for extracting plain text from proprietary formats (PDF, DOCX, ODT, H5). Optimized for preparing context for LLMs and Vector Databases.
An end-to-end Data Science pipeline for analyzing mathematical misconceptions, featuring automated ETL, MySQL integration, and feature engineering for LLM training. Dockerized & CI/CD enabled.
pdf text extraction utility
MDify is a document-to-Markdown conversion library for extracting structured content from complex PDFs and document images, including tables, charts, and scanned documents.
Preflight checks for document extraction pipelines — validate, render, and screen PDFs before they reach your LLM. Pure-Python wheel, in-memory only.
100% client-side privacy-first converter that transforms docs, spreadsheets, code & data into token-optimized .toon files for LLMs. Powered by MarkItDownJS with RAG-ready chunking, semantic optimization & gzip packing — cutting tokens 15-40%. Zero server. Zero telemetry.
To associate your repository with the llm-preprocessing topic, visit your repo's landing page and select "manage topics."