📚A curated list of Awesome LLM/VLM Inference Papers with Codes: Flash-Attention, Paged-Attention, WINT8/4, Parallelism, etc.🎉
-
Updated
Jun 23, 2026 - Python
📚A curated list of Awesome LLM/VLM Inference Papers with Codes: Flash-Attention, Paged-Attention, WINT8/4, Parallelism, etc.🎉
Light-field imaging application for plenoptic cameras
[ICLR 2025] Palu: Compressing KV-Cache with Low-Rank Projection
xKV: Cross-Layer SVD for KV-Cache Compression [ICML 2026]
Decoding Attention is specially optimized for MHA, MQA, GQA and MLA using CUDA core for the decoding stage of LLM inference.
Light field geometry estimator for plenoptic cameras
An efficient and scalable attention module designed to reduce memory usage and improve inference speed in large language models. Designed and implemented the Multi-Head Latent Attention (MLA) module as a drop-in replacement for traditional multi-head attention (MHA) in large language models.
🍺 CLI for quickly generating citations for websites and books
Exam-focused study notes for the AWS Certified Machine Learning - Associate (MLA-C01) certification exam.
Tiny-MoE is a lightweight Mixture-of-Experts language model built entirely from scratch in native PyTorch and trained end-to-end on Kaggle using free 2× NVIDIA T4 GPUs. The project implements modern LLM techniques—including MLA, RoPE, YaRN, streaming pre-training, and efficient inference—without relying on existing model implementations.
Official repo for AMD hybrid models training and inference workflow
A production-grade LLM architecture built from scratch in PyTorch. Features Multi-Head Latent Attention (MLA), Mixture of Experts (MoE), GRPO alignment, and a complete 31-part educational course.
Make your BibTeX perfect. Auto-clean entries, unify conference names (e.g., CVPR, NeurIPS), and generate citation keys for LaTeX & Word.
Will this LLM fit your GPU or Mac? npx fitllm — accurate memory math for MLA/sliding-window/hybrid/MoE architectures that naive VRAM calculators get wrong by up to 18x. Single file, zero deps, conformance-vector tested. MIT.
In this section, predicting the energy efficiency of buildings with machine learning algorithms.
Provided is a Google Apps Script that's soul purpose is to help make MLA writing easier
Add a description, image, and links to the mla topic page so that developers can more easily learn about it.
To associate your repository with the mla topic, visit your repo's landing page and select "manage topics."