Building autonomous AI systems & high-performance data pipelines.
I'm Aditya — a Data Scientist and ML Engineer currently finishing my M.S. in Data Science at Stevens Institute of Technology in Hoboken, NJ. I got into this field because I genuinely enjoy the process of taking messy, real-world problems and turning them into something a model can reason about.
Most of my work revolves around LLM-driven autonomous systems and scalable data pipelines. I've built things like multi-agent RAG systems that orchestrate specialized LLM workers, browser extensions that have to fight through Shadow DOM and framework-level restrictions to automate form filling, and speech-to-speech translation pipelines where every millisecond of latency matters. I care a lot about writing code that actually ships — not just notebooks that run once.
Outside of the typical ML stack, I spend time thinking about AI safety and alignment. I find the challenge of getting deterministic, trustworthy behavior out of probabilistic models genuinely interesting, not just as an academic question but as something that matters the moment you put a model in front of real users. I'm also an AWS Certified AI Practitioner and ML Engineer, which keeps me grounded in how these systems run at scale in production.
| Project | Description |
|---|---|
| Termnova · live demo | Production-grade AI contract intelligence platform with hybrid RAG (dense vector + BM25 reranking), LangGraph multi-agent orchestration, OpenTelemetry distributed tracing, Celery async queues, and hallucination guardrails. |
| Cadence | Open-source CI intelligence for GitHub Actions — analyzes build history over the plain GitHub API, quantifies wasted runtime and dollar costs, and automatically opens the PR that fixes it with zero configuration. |
| Multimodal Fraud Detector · live demo | Multi-agent AI pipeline detecting generative-AI fraud in insurance claims across images, PDFs, and video. Employs a "Jury System" of Qwen-VL forensic vision analysis paired with DeepSeek-R1 / Qwen / GLM critic agents with majority voting. |
| QuantServe | Hardware-aware LLM deployment optimizer & inference systems engine — surrogate latency/VRAM models predicting TTFT and TPOT, custom Triton fused W4A16 dequant kernels, and automated Docker / Kubernetes manifest export. |
| Supplier Intelligence Platform | Real-time supplier intelligence & quality monitoring engine — agentic AI root-cause analysis, edge computer vision defect detection (YOLOv8), and distributed Kafka/Airflow data streaming pipelines. |
Landing fixes and features in ML infrastructure, compilers, agent frameworks, and evaluation tooling.
| Repo | Stars | PR | Impact |
|---|---|---|---|
| great_expectations | Resolved 36 mypy type-check errors across 8 excluded integration test patterns. | ||
| Arize Phoenix | Fixed playground/evaluator clients silently dropping Anthropic & Bedrock tool-choice and strict configurations prior to dispatch. |
||
| kornia | Revived LoFTR CPU accuracy tests; fixed PatchMix bounding boxes & same_on_batch pairing (6 merged PRs). | ||
| sqlfluff | Added full AST grammar for Snowflake CREATE/ALTER/DROP ALERT DDL; fixed reflow alignment for leading-comma T-SQL. |
||
| NVIDIA numba-cuda-mlir | Fixed float-to-bool conversion in MLIR lowering pipeline by comparing against zero instead of raw truncations. | ||
| vllm | Use original context length for mRoPE YaRN correction range | ||
| CrewAI | Explicitly fails when expected evaluation metric has no score instead of silently propagating corrupted agent benchmark results. | ||
| Stanford DSPy | Rejects reserved trajectory as an output field in ReAct — surfaces collision at construction instead of mid-run after billed LM calls. |
||
| Microsoft ONNX Runtime | Reconnected producer edge in the graph optimizer when DivMulFusion substitutes Mul's input. |
||
| MLC-AI xgrammar | Resolved JSON Schema references through array indices for grammar-guided LLM structured output generation. |
| Languages |
|
| ML & AI |
|
| Frameworks |
|
| Data |
|
| Infrastructure |
|
Engineered an enterprise-grade contract analysis platform utilizing LangGraph multi-agent orchestration and Hybrid RAG (dense vector embeddings combined with sparse BM25 reranking). Implemented asynchronous ingestion queues with Celery, end-to-end distributed tracing via OpenTelemetry, and strict quantitative RAG evaluation metrics (faithfulness, context precision, hallucination scoring) to audit complex legal agreements with deterministic precision.
Won 1st place in a Databricks Hackathon by designing a Vision-Critic multi-model jury architecture for insurance claim fraud detection. Visual forensic analysis (OpenCV keyframe extraction and metadata artifact detection) was strictly decoupled from logical LLM deduction to eliminate decision contamination. Built weighted heuristic risk-scoring models that evaluate cross-modal signals across video, images, and transaction metadata.
Architected a concurrent, low-latency audio translation pipeline using a producer-consumer threaded design across Whisper (ASR), MarianMT (Neural Translation), and Meta MMS (TTS). Achieved <3ms Voice Activity Detection latency with Silero VAD, sub-3s end-to-end latency, and dynamic gain-normalization filters to eliminate background static.
Developed a distributed medical computer vision pipeline on Google Cloud Platform (GCP) leveraging PySpark and Spark ML for large-scale fundus image preprocessing. Applied PyTorch deep neural networks with custom CLAHE (Contrast Limited Adaptive Histogram Equalization) contrast-enhancement algorithms to classify diabetic retinopathy severity while preserving critical microvascular pathology.





