Turn Repos & Papers into Skills for Autonomous ML Research
An open library of 5,000+ verified, executable skills distilled from 1,000 ML repositories — and the agent that builds them.
English | 简体中文
flowchart LR
subgraph SRC["📚 Sources"]
direction TB
s0["🌐 Tens of thousands<br>of ML repos"] == "curate" ==> s1["⭐ 1,000 top repos<br>+ 📄 papers · ✍️ blogs"]
end
subgraph CRE["🤖 DisCo Creator Agent"]
direction TB
c1["🔍 Discover<br><i>capabilities</i>"] --> c2["🧠 Distill<br><i>into skills</i>"] --> c3["🧪 Verify<br><i>by execution</i>"]
c3 -. refine .-> c2
end
subgraph LIB["🧠 AREX Skill Library"]
direction TB
l1["🧭 One router<br><i>routes any ML task</i>"] ~~~ l2["📖 5,000+ verified skills<br><i>20 areas · 178 families</i>"] ~~~ l3["🛠 Task-oriented skills<br><i>built per task</i>"]
end
subgraph FIN["🧑💻 Your Agent — unchanged"]
direction TB
u["Claude Code · Codex · DisCo<br><i>loads only what<br>the task needs</i>"] --> r["🔬 <b>Autonomous<br>ML research</b><br>🔓 new abilities<br>🏆 better results<br>⚡ fewer tokens"]
end
SRC ==> CRE ==> LIB ==> FIN
Same agent. Same budget. 2.3× the wins. MLE-bench sets agents loose on 75 Kaggle ML competitions. With AREX skills, vanilla Codex goes from winning medals in 31% of them to 73% — outscoring every public leaderboard entry. Skills, not agent engineering.
- News
- Why AREX-Skill
- From Knowledge to Skills
- What Is an AREX Skill
- How AREX-Skill Is Built: DisCo
- Works With Your Coding Agent
- Library Scale
- Skill Gallery
- Do Skills Make Agents Better Researchers
- Quick Start
- The Bigger Vision
- Contributing
- Documentation
- Acknowledgement
- License
- Citation
- 2026-08-27: The library scales to 1,000 repositories and 5,000+ skills, with a rebuilt router covering every repo. Technical report coming soon, with full benchmark results on MLE-bench, PaperBench, Frontier-CS, and PassNet.
- 2026-08-03: AREX-Skill launches with DisCo's Creator and Researcher workflows and the initial AREX-Skill Library release: 1,000+ operating skills for 170 widely used repositories.
Research knowledge is everywhere. Agents still can't use it well.
| 📄 Papers | 💻 Codebases | 🤖 Agents |
|---|---|---|
| Explain why things work | Contain working implementations | Still have to rediscover everything |
Papers, repositories, and blogs hold nearly all of the field's know-how — but they are written for human readers. They are loose, heterogeneous, and expose no interface an agent can call. So on every task, an agent burns its context window and execution budget searching, reading, and trial-and-erroring its way back to knowledge the field already wrote down.
We call the missing layer operating knowledge: the know-how that separates knowing a method from making it work. AREX-Skill pre-compiles it, once, into skills.
We don't summarize repositories. We compile them into skills agents can execute.
flowchart TB
subgraph D["📚 Descriptive Knowledge — written for humans"]
direction LR
d1["📄 <b>Papers</b><br><i>methods & why they work —<br>but no runnable path</i>"] ~~~ d2["💻 <b>Repos</b><br><i>working code —<br>but usage stays implicit</i>"] ~~~ d3["✍️ <b>Blogs</b><br><i>tricks & pitfalls —<br>scattered, unverified</i>"]
end
D == "⚗️ <b>Skill Distillation</b> — extract · operationalize · verify" ==> O
subgraph O["🛠 Operational Knowledge — built for agents"]
direction LR
o1["🎯 <b>When to use</b><br><i>applicability conditions<br>the router can match</i>"] ~~~ o2["📋 <b>How to use</b><br><i>step-by-step workflows<br>with expected behavior</i>"] ~~~ o3["▶️ <b>What to run</b><br><i>commands, scripts,<br>ready-made tools</i>"]
o4["✅ <b>How to validate</b><br><i>checks & expected<br>observations</i>"] ~~~ o5["🚑 <b>How to recover</b><br><i>known failures with<br>fixes attached</i>"] ~~~ o6["📎 <b>What to trust</b><br><i>evidence linked back<br>to the source</i>"]
end
The output is not a summary and not a RAG index. It is Knowledge → Capability: a skill declares its use conditions, execution behavior, supporting evidence, validation steps, and failure handling — everything an agent needs to act, verify, and recover without re-deriving the source.
An AREX Skill is a self-contained, agent-readable unit of operating knowledge.
It follows the open Agent Skills
format and is organized around SKILL.md, with optional supporting resources:
AREX Skill
│
├── 🧭 SKILL.md: Entry Point / Router
│ ├── applicability and scope
│ ├── task routing
│ ├── operating workflows
│ └── validation and troubleshooting
│
├── 📖 references/: Focused Instructions
│ ├── installation and configuration
│ ├── detailed workflows
│ └── troubleshooting and provenance
│
└── 🛠 scripts/: Executable Helpers
├── diagnostics and smoke tests
├── workflow utilities
└── compatibility and reusable helpers
SKILL.md is the entry point for the skill's scope, operating instructions,
and—when applicable—task routing. references/ and scripts/ are optional:
references provide focused guidance and provenance, while scripts provide
executable helpers for diagnostics, workflows, and repeatable checks.
When a repository exposes several workflows, its skills can be organized as a repository skill graph. A root skill routes the task to focused sub-skills:
Repository Skill Graph
│
├── root SKILL.md: entry point and router
└── sub-skills/
├── inference/SKILL.md
├── training/SKILL.md
└── evaluation/SKILL.md
Each sub-skill is an individual AREX Skill and may have its own references and scripts. The links from the root to its sub-skills form a skill graph—an AREX-Skill extension that supports progressive disclosure, so the agent reads only the branch required by the task.
The AREX-Skill Library is built through DisCo's discover, distill, and verify workflow.
flowchart LR
S["📦 repo · paper · blog"] --> A["🔍 <b>Discover</b><br><i>map what the source<br>can actually do</i>"]
A --> B["🧠 <b>Distill</b><br><i>write skills with evidence,<br>checks & recovery paths</i>"]
B --> C["🧪 <b>Validate</b><br><i>execute examples & tests<br>in a real environment</i>"]
C == "✅ passed" ==> L["📚 <b>Ship</b><br><i>into the library</i>"]
C -- "❌ failed" --> E["🔁 <b>Evolve</b><br><i>repair & refine</i>"]
E --> B
An ordinary repo-to-doc tool stops at Repo → Documentation. DisCo's Creator
agent runs a full experimental loop — evidence-backed exploration, skill-graph
generation, then verification with refinement: generated checks and native
examples are executed, failures are repaired, and the loop repeats until the
graph passes or the budget is spent. Skills ship only after they survive their
own tests.
No new agent. No new workflow.
flowchart LR
S["🧠 <b>AREX Skills</b><br><i>plain SKILL.md graphs<br>+ one library router</i>"] ==> G
subgraph G["your existing agents — workflow unchanged"]
direction LR
A["<b>Claude Code</b><br><i>drop into skills dir</i>"] ~~~ B["<b>Codex</b><br><i>our benchmark harness</i>"] ~~~ C["<b>DisCo</b><br><i>bundled CLI:<br>install · route · update</i>"]
end
Skills are plain SKILL.md graphs in the emerging agent-skills format. Drop
them into the coding agent you already use — no proprietary runtime, no new
research platform to migrate to. The bundled DisCo CLI manages
installation, routing, and updates, and our benchmark results below use
unmodified Codex as the harness.
| widely used ML repositories |
autonomously distilled & verified skills |
research areas, 178 task families |
Sources: GitHub · Papers · Technical Blogs
This is not a demo. It is the first scale point of a growing Machine Learning Research Skill Library — from training infrastructure, LLM alignment, and inference serving to robotics, genomics, and scientific computing. The catalog lists every graph with its upstream repository, source commit, and coverage.
What a skill looks like in practice — source → skill → one prompt:
| Source | Skill | Ask your agent | |
|---|---|---|---|
| 🔍 | FAISS | Vector search & index composition | "Optimize this FAISS index for lower latency at recall ≥ 0.95." |
| ⚡ | vLLM | High-throughput LLM serving | "Benchmark vLLM vs SGLang on this model and report verified throughput." |
| 🧠 | Unsloth | Efficient LLM fine-tuning | "Fine-tune Llama on this dataset within 24 GB VRAM." |
| 🔥 | Diffusers | Diffusion training & inference | "Train a LoRA for this style and validate outputs." |
| 🦾 | LeRobot | Robot learning workflows | "Train and evaluate an ACT policy on this manipulation dataset." |
| 🧬 | AlphaFold2 | Protein structure prediction | "Fold these sequences and check confidence metrics." |
Every graph follows the same contract: a routed entry skill, focused sub-skills for real workflows (data, training, evaluation, serving, troubleshooting), and validation steps the agent can actually run.
We hold everything fixed — Codex harness, GPT-5.5 (xhigh) backbone, same execution budget — and change exactly one thing: whether the agent has AREX-distilled skills.
MLE-bench (medal rate across 75 Kaggle competitions)
without skills ███████░░░░░░░░░░░░░░░░ 31.1%
with skills █████████████████░░░░░░ 72.9% (+134% relative)
PaperBench (replication score, 20 papers)
without skills ███████░░░░░░░░░░░░░░░░ 29.5
with skills █████████░░░░░░░░░░░░░░ 39.6 (+34% relative)
| Benchmark | Metric | Codex | Codex + AREX-Skill | Δ |
|---|---|---|---|---|
| MLE-bench (full, 75 tasks) | Medal rate (Any Medal) % | 31.11 | 72.89 | +41.78 |
| PaperBench (full, 20 papers) | Replication score | 29.45 | 39.59 | +10.14 |
| Frontier-CS (Agent Track, 188 tasks) | Score | 70.63 | 77.14 | +6.51 |
| PassNet (200 samples) | AS Score | 1.343 | 1.531 | +14.0% |
Highlights from the technical report (release coming soon):
- Beyond agent engineering. Vanilla Codex + skills tops the strongest public MLE-bench entries (72.89 vs 64.44) — no custom harness, no modified control loop, only distilled operating knowledge.
- The advantage grows with difficulty. On MLE-bench High-tier tasks the score rises from 13.3% to 62.2% (4.7×); skills matter most exactly where unguided trial-and-error is most expensive.
- Recovery, not just polish. The largest gains land on tasks where the
no-skill agent nearly fails (PaperBench
rice: 7.9 → 48.5; Frontier-CS tasks scoring <50 gain +26.6 on average). - Efficiency, not brute force. On Frontier-CS, Codex + AREX-Skill Pareto-dominates the leaderboard's Claude Code entries on score, tokens, steps, and tool calls, using ~3× fewer tokens.
Without AREX With AREX
──────────── ─────────
Search the web Route to skill
Inspect the repo Execute known workflow
Guess an approach Validate against checks
Debug from scratch Optimize the target metric
Retry, repeat…
Benchmarks show that performance improves; the operating pattern shows why: the agent enters a productive region of the solution space early and spends its budget on the choices that move the metric. See a full Creator session building a FlagEmbedding graph and a Researcher session applying Gymnasium + Stable-Baselines3 skills to an auditable RL experiment.
Three steps to a skill-powered research agent:
# 1. Install the DisCo CLI (Node.js >= 22.19)
npm install -g @auto-ml-skills/disco
# 2. Install the skill library (1,000 repos + router)
disco repo-skills install
# 3. Research with skills
disco -p "Use the installed skills to benchmark vLLM and SGLang \
on this machine and report verified throughput for each."Configure a model provider on first run with /login or environment variables
(ANTHROPIC_API_KEY, OPENAI_API_KEY, GEMINI_API_KEY, …).
Create your own skills (Creator mode)
DisCo's Creator mode distills new skill graphs from any repository or paper, then verifies them before import:
git clone https://github.com/FlagOpen/FlagEmbedding.git
disco --agent-mode creator -p \
"/skill:distill-ml-knowledge Create and verify a repository skill graph \
for ./FlagEmbedding covering embedding inference and evaluation."Researcher is the default mode; switch with --agent-mode creator|researcher
or /creator · /researcher in the UI. See
DisCo Workflows for the 15 bundled Creator meta
skills, verification gates, and maintenance workflows.
Use the skills in another agent, manage the collection, build from source
- Other agents: skills are standard
SKILL.mdgraphs; see Meta Skills For Other Agents for Claude Code / Codex installation. - Manage:
disco repo-skills status | update, router toggle withdisco repo-skills router disable|enable. - Manual install: copy
skills/repositories/repo-skillsandskills/repositories/repo-skills-routerinto~/.disco/agent/skills/repositories/. - From source:
bash scripts/build-from-source-link.shafter cloning.
The full Installation Guide covers provider setup, update/backup semantics, router behavior, and every fallback path.
From repositories to skills. From skills to autonomous research.
Today's research knowledge is written for humans. AREX turns it into operating knowledge for AI researchers.
flowchart LR
A["📦 <b>Repos + Papers</b><br><i>written for humans</i>"] --> B["🧠 <b>Research Skills</b><br><i>distilled once, verified</i>"]
B --> C["🌐 <b>Skill Ecosystem</b><br><i>shared & inherited</i>"]
C --> D["🤖 <b>Autonomous Agents</b><br><i>start where the field left off</i>"]
D --> E["⚙️ <b>Automated<br>ML R&D</b>"]
Every skill distilled once is inherited by every agent afterward. As the library grows — more repositories, more papers, more task families — each research task starts a little further from zero. We believe ML research knowledge should exist not only as papers and repos, but as skills that AI researchers can directly call.
Naming note: AREX-Skill is the project and the published AREX-Skill Library. DisCo is the bundled skill-powered CLI/runtime that creates skills (Creator) and researches with them (Researcher).
We welcome three kinds of contributions — new repo skills, refreshes of existing skills, and DisCo CLI improvements. Skill PRs should include provenance (model, source commit, verification steps); see CONTRIBUTING.md for the checklist and Contributing docs for the repo-skill layout and router update workflow.
| Page | Description |
|---|---|
| Installation Guide | Full CLI and skill-collection installation, provider setup, router toggle, manual fallback. |
| DisCo Workflows | Modes, sessions, Researcher execution, Creator construction, deployment scopes. |
| AREX-Skill Library | Library model, collection layout, installation. |
| Imported Repo Skills Catalog | Every published graph with upstream baselines. |
| Repository Catalog | Human-readable area -> family inventory of all published repository skills. |
| Architecture | Repository layers, routing, authoring pipelines, deployment scopes. |
| Examples | Sanitized end-to-end Creator and Researcher sessions. |
| Bundled Skills Reference | Creator meta-skill contracts and artifact layouts. |
| DisCo CLI README | CLI usage, runtime skill routing, packages. |
DisCo's CLI and agent runtime are built on the foundation of earendil-works/pi, an open-source AI agent toolkit with a unified LLM API, agent loop, terminal UI, and coding-agent CLI.
AREX-Skill is also made possible by the open-source community on GitHub. The repository skills in this library build on high-quality ML, agent, data, bio/chemistry, vision, and infrastructure projects released by researchers and engineers around the world. We are grateful to everyone who makes that work available for the community to use and build on.
Unless a file or component states otherwise, repository-level AREX-Skill
materials are released under the Apache License 2.0. The skills published in
the library are licensed separately: before using, copying, modifying, or
redistributing a skill, inspect the license metadata field in that skill's
SKILL.md. That license is authoritative for the individual skill, may differ
from this repository's Apache-2.0 license, and may include additional terms or
restrictions. Users are responsible for reviewing and complying with the terms
of every individual skill they use.
The standalone DisCo npm package under cli/ is distributed under its
own MIT License, with upstream attribution in
cli/THIRD_PARTY_NOTICES.md.
TBA