Pinned Loading
-
llm-router
llm-router PublicCache-aware router for OpenAI-compatible LLM servers, in Go. Per-worker radix trees route each request to the worker holding its KV prefix. Validated on 4x A100 + vLLM and Apple Silicon + llama.cpp.
-
distill-sql
distill-sql PublicOn-device text-to-SQL distilled from GPT-4o-mini into Qwen2.5 (0.5B → 3B locally on M1 via mlx-lm LoRA, 7B+ on cloud A100). 847 MB at 62.5% on Spider dev; 3B variant hits 72.6%, 7B variant hits 75.0%.
Python 1
-
folio-writer
folio-writer PublicMulti-agent article generator: Spring AI Alibaba StateGraph, token streaming over SSE, parallel image generation across six providers, Vue 3 frontend.
Java 2
-
gpu-k8s-operator
gpu-k8s-operator PublicKubernetes operator for rolling-window GPU-hour budgets. Stateless accounting (recomputes from API state every reconcile, 0.996 accuracy after operator kill). Enforcement via eviction, pause, or al…
Go
If the problem persists, check the GitHub status page or contact support.


