An advanced, production-ready Full-Stack AI Application that implements a Multi-Agent workflow using LangGraph, LangChain, FastAPI, ChromaDB, and React (Vite) to parse, index, analyze, compare and discuss research papers in natural language.
- Ingestion Pipeline: Upload multiple PDF research papers. Text is extracted, parsed for tables, and automatically run through Tesseract OCR when scanned layouts are detected.
- Dynamic Chunking & Embeddings: Documents are recursively split by characters, preserving page numbers and source metadata. Embeddings are created locally using
sentence-transformers/all-MiniLM-L6-v2. - ChromaDB Vector Store: PERSISTENT storage client allowing semantic, context-aware similarity retrieval.
- 15 Specialized LangGraph Agents:
- Planner Agent: Route queries to the appropriate sub-agents.
- PDF Parsing Agent: Extract text and metadata.
- OCR Agent: Clean extract scanned pages via Tesseract.
- Chunking Agent: Intelligently segment pages.
- Embedding Agent: Index chunks in ChromaDB.
- Retrieval Agent: Fetch context matched with query.
- Summary Agent: Generate abstract summaries, executive insights, novelty and key contributions.
- Methodology Agent: Identify math equations, algorithms, and pipelines.
- Experiment Agent: Extract datasets, hyperparams, and training steps.
- Result Agent: Explain tables, graphs, precision/recall metrics.
- Comparison Agent: Correlate similarities/differences of multiple papers.
- Citation Agent: Formats IEEE, APA, MLA, and BibTeX.
- Research Gap Agent: Synthesize missing work and future directions.
- Report Agent: Synthesize structured Markdown summaries.
- Chat Agent: Conversational RAG chatbot referencing page citations and history memory.
- Replaceable LLM Provider: Toggle between Google Gemini, OpenAI, Claude, and Groq by changing only
settings.pyor.envfiles.
- Frontend: React (Vite), Tailwind CSS (v3), Axios, React Router, React Markdown, Lucide Icons, Canvas Confetti.
- Backend: FastAPI, Python 3.12, Uvicorn, LangGraph, LangChain, SQLite.
- OCR: Tesseract OCR.
- Vector DB: ChromaDB.
/ (Project Root)
├── backend/
│ ├── app/
│ │ ├── api/ # FastAPI Routes (upload, chat, papers, analysis)
│ │ ├── config/ # Environment configs (settings.py)
│ │ ├── database/ # ChromaDB Client & SQLite connection
│ │ ├── embeddings/ # Local embedding model loader
│ │ ├── models/ # State & Database Schemas
│ │ ├── parser/ # PDF parsing & OCR trigger logic
│ │ ├── ocr/ # Tesseract OCR Wrapper
│ │ ├── rag/ # RAG pipeline logic
│ │ ├── agents/ # The 15 agents (planner, summary, analysis_agents, etc.)
│ │ ├── graph/ # LangGraph State & Workflow definition (workflow.py)
│ │ └── main.py # FastAPI Entrypoint
│ └── requirements.txt
├── frontend/
│ ├── src/
│ │ ├── components/ # Chat, Sidebar
│ │ ├── pages/ # Dashboard, Library, Upload, Viewer, Comparison, Settings
│ │ ├── services/ # Axios api wrapper
│ │ ├── context/ # App context states
│ │ ├── App.jsx
│ │ └── main.jsx
│ ├── tailwind.config.js
│ ├── index.html
│ └── netlify.toml # Netlify deployment configuration
└── netlify.toml # Monorepo Netlify build configuration
- Python 3.12+
- Node.js 18+
- Tesseract OCR (Install and add to PATH environment variable to enable OCR capabilities for scanned PDFs)
- Navigate to the backend directory:
cd backend - Create and activate a virtual environment:
python -m venv venv # On macOS/Linux: source venv/bin/activate # On Windows: venv\Scripts\activate
- Install the required dependencies:
pip install -r requirements.txt
- Start the FastAPI server:
The backend API documentation will be available at
uvicorn app.main:app --reload --port 8000
http://localhost:8000/docs.
- Navigate to the frontend directory:
cd ../frontend - Install the package dependencies:
npm install
- Run the development server:
The application UI will be running at
npm run dev
http://localhost:5173.
Create a .env file in the backend/ directory to configure API keys:
LLM_PROVIDER=google
GEMINI_API_KEY=your_gemini_api_key_here
OPENAI_API_KEY=your_openai_api_key_here
GROQ_API_KEY=your_groq_api_key_hereThe frontend of this project is prepared for deployment to Netlify.
This repository includes a root netlify.toml file. When you link this project to Netlify, it automatically detects the configuration and sets:
- Base directory:
frontend - Build command:
npm run build - Publish directory:
dist
To connect the deployed frontend with your backend API:
- In the Netlify dashboard, go to Site settings > Environment variables.
- Add a new variable:
- Key:
VITE_API_BASE_URL - Value:
https://your-backend-api-url.com/api/v1(The production URL where your FastAPI backend is hosted).
- Key:
Note
Since the backend uses ChromaDB (local persistence) and SQLite, it should be deployed on a platform like Render, Railway, Fly.io, or AWS with persistent storage volumes configured for backend/data/.