VasukiSquare is an AI-powered ebook generation engine written in Python. It researches, plans, writes, designs, renders, and compiles complete, publication-ready physical A4 PDF ebooks from a topic or editorial brief.
VasukiSquare solves the problem of unstructured, repetitive, and poorly formatted AI-generated documents. Instead of dumping raw markdown into a single prompt, VasukiSquare uses a multi-stage reasoning pipeline that performs web research, enforces fact-checking and domain citation hierarchies, structures cohesive chapters, designs high-contrast cover art, applies mathematical visual design tokens, and renders pixel-perfect physical A4 PDFs with zero overflow.
- Infers Intent & Audience: Determines whether the topic is technical or non-technical, identifies target reader depth, and extracts required themes and elements.
- Conducts Multi-Perspective Research: Expands the topic into search queries across documentation, specifications, and real-world patterns using swappable search providers (DuckDuckGo, SearXNG, Tavily, Serper, Brave, or Wikipedia).
- Plans Editorial Structure & Page Budgets: Outlines chapters, assigns section visual anchors (code blocks, comparison tables, callout tips, checklists), and budgets physical A4 page counts.
- Designs Custom Cover Artwork: Generates 1600 × 2560 canvas cover art using one of 6 layout styles, procedural math patterns, and WCAG-compliant contrast validation.
- Authors Structured Page Content: Writes structured component blocks without markdown artifacts, applying alternating dark/light chapter themes and 6 distinct chapter opener templates.
- Audits & Repairs Physical Layout: Evaluates A4 content density, dynamically splits overflowing sections, and recalculates dynamic Table of Contents page numbers.
- Compiles Publication Artifacts: Renders HTML/CSS templates and exports physical A4 PDFs using Playwright headless Chromium.
- Topic & Brief Driven Generation: Generate complete books from a single topic string (
--topic) or a detailed editorial brief (--prompt/--prompt-file). - Physical A4 PDF Compilation: Exact ISO A4 dimensions (210mm × 297mm) rendered via Playwright Chromium with zero vertical overflow.
- Dual LLM Provider Architecture:
- Groq Cloud: High-speed inference (200–500+ tokens/sec) with multi-key rotation pools, model failover pools, and cooldown management.
- Ollama Local: 100% private, offline generation with zero external cloud API fees.
- Swappable Web Research Backends: Integrated support for DuckDuckGo (free zero-key search), SearXNG (local metasearch), Tavily, Serper, Brave Search, and offline mock modes.
- Visual Design System: Strict mathematical color tokens, Lucide SVG icons, responsive typography stacks, and alternating chapter themes (odd = dark mode, even = light mode).
- 6 Distinct Chapter Opener Styles: Unique opener templates (
minimal_centered,left_accent_banner,split_contrast,editorial_classic,technical_blueprint,icon_heroic). - High-Resolution Cover Engine: 1600 × 2560 source canvas generator supporting 6 cover styles (
editorial_minimal,split_hero,geometric_accent,technical_blueprint,swiss_bold,minimal_monochrome) with automated contrast repair. - White-Label Publisher Branding: Complete customization of author names, publishing imprints, copyright notices, edition names, and website URLs via
config.json. - Checkpoint & Resume Support: Granular JSON stage and page checkpointing allows interrupted runs to resume seamlessly with
--resume. - 15-Point Content Quality Audit: Automated validation checks structure, density, semantic validity, and code formatting before final assembly.
- Production MongoDB Architecture: Canonical 4-collection publication store (
books,pages,covers,ads) with deterministic URL-safe slugs, featured slots 1..5, native ad inventory, atomic counters, and headless Next.js query support.
Every generated publication produces a dedicated output folder containing:
output/
├── book.pdf # Physical A4 compiled PDF ready for reading or distribution
├── book.html # Assembled standalone HTML document with embedded CSS
├── cover.html # Standalone high-resolution cover artwork
├── book_manifest.json # Metadata, chapter manifest, theme tokens, and page counts
├── book_plan.json # Editorial plan, page budgeting, and section visual anchors
├── research.json # Deduplicated and ranked research sources
├── preflight_report.json # Layout geometry and preflight audit report
├── generation_metrics.json # Telemetry (token counts, duration, latency, search queries)
├── pages/ # Standalone HTML files for each individual A4 page
│ ├── page_001.html # Cover page
│ ├── page_002.html # Imprint & copyright page
│ ├── ... # Content pages & chapter openers
│ └── page_040.html # Backmatter / References page
└── checkpoints/ # Stage checkpoints for crash recovery
For practical generation examples, see the Examples Directory:
- Basic Generation Example
- Technical Programming Book Example
- Non-Technical / Productivity Book Example
User Topic / Prompt
│
▼
Stage 1: Intent Inference (EditorialPlannerAgent)
│
▼
Stage 2: Deep Research (ResearchService & WebSearchTool)
│
▼
Stage 3: Editorial & Chapter Planning (BookPlan & Page Budget)
│
▼
Stage 4: Cover Planning & Design (CoverPlanner & ContrastValidator)
│
▼
Stage 5: Page Authoring & Repair (PageWriterAgent & PageRepairEngine)
│
▼
Stage 6: HTML Assembly, Preflight Audit & PDF Export (HtmlPageRenderer & PdfRenderer)
│
▼
Stage 7: Canonical Database Persistence (Optional MongoDB Linked Graph)
│
▼
Physical A4 PDF (book.pdf) & Canonical Web Store (MongoDB)
For full details on boundaries between deterministic Python rules, LLM reasoning, and rendering, see System Architecture.
- Operating System: Windows 10/11, macOS 12+, or modern Linux (Ubuntu 20.04+, Debian 11+, Fedora 38+).
- Python: Python 3.11, 3.12, or 3.13 (64-bit).
- Browser Runtime: Playwright Chromium (installed via
playwright install chromium). - AI Provider (One of the following):
- Groq API Key (for fast cloud generation; free tier available at console.groq.com), OR
- Ollama running locally with
qwen2.5:7b-instructorllama3.1:8b(for 100% offline, zero-cost generation).
- MongoDB (Optional): Only required if you wish to persist records to a MongoDB instance. Book generation and PDF export function completely without MongoDB.
-
Clone or extract the repository:
git clone https://github.com/UltronTheAI/VasukiSquare.git cd VasukiSquare -
Create and activate a virtual environment:
Linux / macOS (Bash / Zsh):
python3 -m venv .venv source .venv/bin/activateWindows (PowerShell):
python -m venv .venv .venv\Scripts\Activate.ps1 -
Install VasukiSquare in editable development mode:
pip install --upgrade pip pip install -e .(To include test and lint tooling:
pip install -e ".[dev]") -
Install the Playwright Chromium browser binary:
playwright install chromium
pip install -r requirements.txt
playwright install chromium-
Copy the environment configuration template:
cp .env.example .env
(On Windows PowerShell:
Copy-Item .env.example .env) -
Open
.envand add your Groq API key:GROQ_API_KEY=gsk_your_groq_api_key_here
(Or set
LLM_PROVIDER=ollamato run with local Ollama) -
Generate your first ebook:
vasukisquare --topic "Modern Distributed Systems: Consensus, Raft, and Gossip Protocols" --pages 30 -
Open
output/book.pdfin your PDF reader oroutput/book.htmlin your web browser!
usage: vasukisquare [-h] --topic TOPIC [--title TITLE] [--prompt PROMPT]
[--prompt-file PROMPT_FILE] [--pages PAGES]
[--output-dir OUTPUT_DIR] [--no-pdf] [--no-db]
[--ollama-model OLLAMA_MODEL]
[--llm-provider {groq,ollama,auto}] [--resume]
| Flag | Type | Default | Description |
|---|---|---|---|
--topic TOPIC |
str |
Required | The core topic or subject of the ebook to generate. |
--title TITLE |
str |
None |
Optional explicit public title (must be ≤ 50 characters). |
--prompt PROMPT |
str |
None |
Editorial brief specifying audience, tone, required concepts. |
--prompt-file FILE |
str |
None |
Path to text file containing editorial brief (mutually exclusive with --prompt). |
--pages PAGES |
int |
60 |
Target physical A4 page count for the complete ebook. |
--output-dir DIR |
str |
./output |
Output directory for rendered PDF, HTML, and JSON artifacts. |
--no-pdf |
flag |
False |
Skip Playwright PDF compilation and only output HTML and JSON. |
--no-db |
flag |
False |
Skip persisting records to MongoDB. |
--ollama-model MODEL |
str |
None |
Specify local Ollama model (e.g. qwen2.5:7b-instruct). |
--llm-provider MODE |
str |
None |
Force LLM provider (groq, ollama, or auto). |
--resume |
flag |
False |
Resume generation from existing stage checkpoints in output directory. |
-h, --help |
flag |
False |
Show CLI help message and exit. |
(You can also run python scripts/generate_book.py [OPTIONS] with identical arguments).
vasukisquare \
--topic "Python 3.12 Fundamentals: From Zero to Object-Oriented Programming" \
--title "Python 3.12 Fundamentals" \
--pages 40 \
--output-dir "./output/python-basics"vasukisquare \
--topic "Building Sustainable Daily Routines: The Psychology of Micro-Habits" \
--title "Atomic Daily Routines" \
--prompt "A practical self-improvement guide for knowledge workers. Focus on actionable exercises, reflection checklists, and a 30-day habit roadmap." \
--pages 35 \
--output-dir "./output/daily-routines"vasukisquare \
--topic "The Engineering Behind Modern Databases: B-Trees, WAL, MVCC and Distributed Storage" \
--title "Modern Database Internals" \
--pages 60 \
--output-dir "./output/database-internals"vasukisquare \
--topic "Network Security and Penetration Testing Fundamentals" \
--llm-provider ollama \
--ollama-model qwen2.5:7b-instruct \
--pages 25 \
--output-dir "./output/network-security"VasukiSquare features an automated, unattended promotion engine that periodically selects published ebooks from MongoDB, writes comprehensive, high-signal technical articles using Groq LLMs, and publishes them to developer platforms (starting with DEV Community / dev.to).
Published Books (MongoDB) ──► Fair Weighted Selector ──► Groq AI Writer ──► MongoDB Audit ──► DEV.to API
-
Fair Weighted Selection: Prioritizes ebooks with fewer historical promotions (
$weight = \frac{1}{1 + \text{promotion_count}}$ ) and strictly enforces configurable cooldown periods (PROMOTION_BOOK_COOLDOWN_HOURS=72). -
High-Signal Technical Writing: Generates educational tutorials teaching core book concepts, avoiding clickbait and spammy marketing while seamlessly linking back to the canonical web reader route (
https://vasukisquare.cc/book/{slug}). -
Two-Phase Idempotency: Generated posts are persisted in MongoDB (
promotion_posts) before attempting publication, preventing duplicate posts upon GitHub Action retries. - Dry-Run Simulation: Test selection, prompt formatting, and LLM output without sending live requests to external publisher APIs.
# Run standard campaign (promotes configured count, default 3 books to DEV)
python scripts/promotion.py
# Simulate promotion (generates and stores posts in MongoDB, skips DEV publication)
python scripts/promotion.py --dry-run
# Promote custom number of books
python scripts/promotion.py --count 5
# Target a specific book by ID or slug
python scripts/promotion.py --book "modern-database-internals" --dry-runVasukiSquare separates infrastructure secrets from publisher branding:
- Environment Variables (
.env/.env.local): Credentials, model pools, search providers, and performance limits. - Publication Metadata (
config.json): Author names, publisher imprints, corporate identities, edition names, and copyright notices.
| Variable | Required? | Default | Description | Example |
|---|---|---|---|---|
LLM_PROVIDER |
Optional | auto |
Active provider (auto, groq, ollama) |
LLM_PROVIDER=auto |
GROQ_API_KEY |
Conditional | None |
Single or comma-separated list of Groq API keys | GROQ_API_KEY=gsk_key1,gsk_key2 |
GROQ_MODEL |
Optional | openai/gpt-oss-120b |
Primary Groq model | GROQ_MODEL=openai/gpt-oss-120b |
GROQ_MODELS |
Optional | None |
Comma-separated Groq model pool for automatic failover | GROQ_MODELS=openai/gpt-oss-120b,llama-3.3-70b-versatile |
OLLAMA_BASE_URL |
Optional | http://localhost:11434 |
Ollama local endpoint | OLLAMA_BASE_URL=http://localhost:11434 |
OLLAMA_MODEL |
Optional | qwen2.5:7b-instruct |
Ollama model name | OLLAMA_MODEL=qwen2.5:7b-instruct |
SEARCH_PROVIDER |
Optional | auto |
Active search backend (auto, duckduckgo, searxng, tavily, serper, brave, mock) |
SEARCH_PROVIDER=auto |
SEARXNG_URL |
Optional | http://localhost:8080 |
Local SearXNG endpoint | SEARXNG_URL=http://localhost:8080 |
TAVILY_API_KEY |
Optional | None |
Tavily search API key | TAVILY_API_KEY=tvly-... |
PDF_OUTPUT_DIR |
Optional | ./output |
Default output folder | PDF_OUTPUT_DIR=./output |
MONGODB_URI |
Optional | mongodb://localhost:27017 |
Optional MongoDB connection URI | MONGODB_URI=mongodb://localhost:27017 |
DEVTO_API_KEY |
Optional | None |
DEV Community API key for automated ebook promotion | DEVTO_API_KEY=... |
PUBLICATION_BASE_URL |
Optional | https://vasukisquare.cc |
Canonical public reader base URL | PUBLICATION_BASE_URL=https://vasukisquare.cc |
For the complete variable catalog, see Configuration Guide.
To customize the author name, publisher imprint, website, and copyright notices on all generated publications, edit config.json at the root of the project:
{
"branding": {
"author_name": "Jane Doe",
"publication_name": "Northstar Books",
"company_name": "Northstar Media LLC",
"engine_name": "VasukiSquare AI Publishing Engine",
"website": "https://northstarbooks.example.com"
},
"edition": {
"name": "FIRST EDITION",
"year": 2026
},
"copyright": {
"holder": "Northstar Media LLC",
"all_rights_reserved": true
},
"book_defaults": {
"language": "English"
}
}VasukiSquare/
├── README.md # Primary product documentation
├── LICENSE # Apache 2.0 open-source software license
├── CHANGELOG.md # Keep a Changelog release history
├── CONTRIBUTING.md # Contributor and developer guidelines
├── SECURITY.md # Security policy and secret safety rules
├── DESIGN.md # Visual design system source of truth
├── TESTS.md # Testing contracts and verification rules
├── pyproject.toml # Canonical packaging configuration and dependencies
├── setup.py # Setuptools compatibility shim
├── requirements.txt # Runtime dependencies export
├── .env.example # Complete environment configuration template
├── .gitignore # Excludes secrets, caches, and build artifacts
├── config.json # Validated publisher branding and copyright metadata
├── src/vasukisquare/ # Engine implementation
│ ├── cli.py # Unified CLI parser and async execution runner
│ ├── agents/ # Reasoning agents (editorial, writer, cover, validator)
│ ├── book/ # Domain models, component blocks, layout enums
│ ├── config/ # Pydantic Settings and AppConfig loaders
│ ├── cover/ # 1600x2560 canvas cover generator & contrast validator
│ ├── database/ # PyMongo connection and document graph repositories
│ ├── design/ # Color tokens, themes, layout engine, Lucide icons
│ ├── llm/ # LLMClient, Groq model/key pools, telemetry metrics
│ ├── pipeline/ # EbookGenerationPipeline orchestrator and state
│ ├── renderer/ # HTML assembler, zero-overflow repair, Playwright PDF exporter
│ ├── research/ # Query planner, source ranking, deduplication
│ ├── templates/ # Jinja2 templates (book.html, base.html, styles.css)
│ └── tools/ # Search tools (DuckDuckGo, SearXNG, Tavily, Wikipedia)
├── scripts/ # Utility and standalone generation scripts
│ ├── generate_book.py # Standalone book generation script
│ ├── generate_demo.py # End-to-end demo verification script
│ └── verify_dom.py # DOM structure inspector
├── tests/ # 234+ automated tests
│ ├── unit/ # Unit tests (config, tokens, agents, pools, models)
│ ├── integration/ # Integration tests (pipeline, research, database)
│ └── rendering/ # Physical A4 rendering & snapshot contract tests
├── examples/ # Example commands and sample configuration files
└── docs/ # Complete documentation suite
VasukiSquare maintains an extensive automated test suite with 100% mock mode support (requiring zero live API keys or active database servers to run):
# Run all tests
pytest -q
# Run unit tests
pytest tests/unit -q
# Run integration tests
pytest tests/integration -q
# Run rendering validation tests
pytest tests/rendering -q| Guide | Description |
|---|---|
| Documentation Hub | Complete index of all technical and user guides |
| Getting Started | Step-by-step onboarding walkthrough for beginners |
| User Guide | Practical generation workflows and options |
| CLI Reference | Exhaustive command-line arguments and flags |
| Automation & GitHub Actions | Production scheduling, GitHub Actions workflows, and MongoDB queue lifecycle |
| Configuration Guide | Environment variables, .env options, and config.json |
| LLM & Search Providers | Groq, Ollama, DuckDuckGo, SearXNG, and commercial search setup |
| Customization Guide | Extending prompts, component blocks, color tokens, and themes |
| System Architecture | Package architecture, Mermaid data-flows, and layer boundaries |
| Generation Pipeline | 7-stage generation lifecycle and checkpoint mechanics |
| Research Engine | Multi-perspective search, ranking, and citation synthesis |
| Editorial Planning | Intent analysis, title rules, and page budgeting |
| Design System | Physical A4 layout contracts, color tokens, and chapter openers |
| Cover Design Engine | 1600 × 2560 canvas, 6 cover styles, and contrast validation |
| Rendering & PDF Export | Playwright Chromium PDF compilation and zero-overflow repair |
| MongoDB Schema | Optional document graph persistence and collection models |
| Troubleshooting | Diagnosis and fixes for common installation and runtime errors |
| FAQ | Frequently asked technical and commercial questions |
| Commercial Distribution | Source-code package details and API key guidance |
| Release Checklist | Pre-distribution quality and packaging verification |
When you purchase the VasukiSquare source code, you receive the full Python codebase to run, customize, and self-host the publishing engine.
- Infrastructure: Customers run the engine on their own machines or cloud infrastructure.
- API Keys: Customers supply their own API keys where applicable (e.g. Groq Cloud, Tavily) or use free local backends (Ollama, DuckDuckGo, SearXNG).
- Never commit
.envor.env.localfiles. The repository.gitignoreexcludes local environment files by default. - If an API key is ever exposed, rotate it immediately in your provider dashboard.
- For vulnerability reporting guidelines, please see SECURITY.md.
VasukiSquare is licensed under the Apache License 2.0. Commercial rights for generated book text depend on your applicable model provider terms and citation attribution rules.
For bugs, questions, and feature discussions, please use the official repository issue tracker: GitHub Issues