Skip to content

About

genAI pep

Resources

Stars

0 stars

Watchers

0 watching

Forks

Latest commit

 

History

2 Commits

Folders and files

Repository files navigation

RAG-Research-Anaysis-Project

Research Paper Analysis Agent (Multi-Agent RAG System)

An advanced, production-ready Full-Stack AI Application that implements a Multi-Agent workflow using LangGraph, LangChain, FastAPI, ChromaDB, and React (Vite) to parse, index, analyze, compare and discuss research papers in natural language.


🚀 Key Features

  • Ingestion Pipeline: Upload multiple PDF research papers. Text is extracted, parsed for tables, and automatically run through Tesseract OCR when scanned layouts are detected.
  • Dynamic Chunking & Embeddings: Documents are recursively split by characters, preserving page numbers and source metadata. Embeddings are created locally using sentence-transformers/all-MiniLM-L6-v2.
  • ChromaDB Vector Store: PERSISTENT storage client allowing semantic, context-aware similarity retrieval.
  • 15 Specialized LangGraph Agents:
    1. Planner Agent: Route queries to the appropriate sub-agents.
    2. PDF Parsing Agent: Extract text and metadata.
    3. OCR Agent: Clean extract scanned pages via Tesseract.
    4. Chunking Agent: Intelligently segment pages.
    5. Embedding Agent: Index chunks in ChromaDB.
    6. Retrieval Agent: Fetch context matched with query.
    7. Summary Agent: Generate abstract summaries, executive insights, novelty and key contributions.
    8. Methodology Agent: Identify math equations, algorithms, and pipelines.
    9. Experiment Agent: Extract datasets, hyperparams, and training steps.
    10. Result Agent: Explain tables, graphs, precision/recall metrics.
    11. Comparison Agent: Correlate similarities/differences of multiple papers.
    12. Citation Agent: Formats IEEE, APA, MLA, and BibTeX.
    13. Research Gap Agent: Synthesize missing work and future directions.
    14. Report Agent: Synthesize structured Markdown summaries.
    15. Chat Agent: Conversational RAG chatbot referencing page citations and history memory.
  • Replaceable LLM Provider: Toggle between Google Gemini, OpenAI, Claude, and Groq by changing only settings.py or .env files.

🛠️ Tech Stack

  • Frontend: React (Vite), Tailwind CSS (v3), Axios, React Router, React Markdown, Lucide Icons, Canvas Confetti.
  • Backend: FastAPI, Python 3.12, Uvicorn, LangGraph, LangChain, SQLite.
  • OCR: Tesseract OCR.
  • Vector DB: ChromaDB.

📂 Project Structure

/ (Project Root)
├── backend/
│   ├── app/
│   │   ├── api/          # FastAPI Routes (upload, chat, papers, analysis)
│   │   ├── config/       # Environment configs (settings.py)
│   │   ├── database/     # ChromaDB Client & SQLite connection
│   │   ├── embeddings/   # Local embedding model loader
│   │   ├── models/       # State & Database Schemas
│   │   ├── parser/       # PDF parsing & OCR trigger logic
│   │   ├── ocr/          # Tesseract OCR Wrapper
│   │   ├── rag/          # RAG pipeline logic
│   │   ├── agents/       # The 15 agents (planner, summary, analysis_agents, etc.)
│   │   ├── graph/        # LangGraph State & Workflow definition (workflow.py)
│   │   └── main.py       # FastAPI Entrypoint
│   └── requirements.txt
├── frontend/
│   ├── src/
│   │   ├── components/   # Chat, Sidebar
│   │   ├── pages/        # Dashboard, Library, Upload, Viewer, Comparison, Settings
│   │   ├── services/     # Axios api wrapper
│   │   ├── context/      # App context states
│   │   ├── App.jsx
│   │   └── main.jsx
│   ├── tailwind.config.js
│   ├── index.html
│   └── netlify.toml      # Netlify deployment configuration
└── netlify.toml          # Monorepo Netlify build configuration

💻 Setup and Installation

Prerequisites


Local Installation

1. Setup Backend

  1. Navigate to the backend directory:
    cd backend
  2. Create and activate a virtual environment:
    python -m venv venv
    # On macOS/Linux:
    source venv/bin/activate
    # On Windows:
    venv\Scripts\activate
  3. Install the required dependencies:
    pip install -r requirements.txt
  4. Start the FastAPI server:
    uvicorn app.main:app --reload --port 8000
    The backend API documentation will be available at http://localhost:8000/docs.

2. Setup Frontend

  1. Navigate to the frontend directory:
    cd ../frontend
  2. Install the package dependencies:
    npm install
  3. Run the development server:
    npm run dev
    The application UI will be running at http://localhost:5173.

⚙️ Configuration (.env)

Create a .env file in the backend/ directory to configure API keys:

LLM_PROVIDER=google
GEMINI_API_KEY=your_gemini_api_key_here
OPENAI_API_KEY=your_openai_api_key_here
GROQ_API_KEY=your_groq_api_key_here

🌐 Deployment to Netlify

The frontend of this project is prepared for deployment to Netlify.

1. Automated Setup (Recommended)

This repository includes a root netlify.toml file. When you link this project to Netlify, it automatically detects the configuration and sets:

  • Base directory: frontend
  • Build command: npm run build
  • Publish directory: dist

2. Environment Variables

To connect the deployed frontend with your backend API:

  1. In the Netlify dashboard, go to Site settings > Environment variables.
  2. Add a new variable:
    • Key: VITE_API_BASE_URL
    • Value: https://your-backend-api-url.com/api/v1 (The production URL where your FastAPI backend is hosted).

Note

Since the backend uses ChromaDB (local persistence) and SQLite, it should be deployed on a platform like Render, Railway, Fly.io, or AWS with persistent storage volumes configured for backend/data/.

About

genAI pep

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages