From c5a511759faeca0eb344425c3731b3a9e0abedac Mon Sep 17 00:00:00 2001 From: Gotham-Zolio <18781106300@163.com> Date: Wed, 16 Sep 2026 19:30:07 -0400 Subject: [PATCH] docs: E11 on the home page, and the pi0 examples that could not run Three changes, all pointing at the same run. The home page gains a section after Verify. The same two processes carry a full-size pi0.5: the env client steps LIBERO, the server answers with actions, FPO trains on the feedback. As a control the unmodified checkpoint scored 99 of 100 on libero_spatial and 185 of 200 on libero_10, against openpi's published 98.8 and 92.4, with the server's episode and step counts equal to the clients'. The reinforcement learning result is negative and the section says so: one FPO iteration took the hardest task from 26 of 50 to 0 of 50, at one iteration of ten, because a second learn step does not fit beside the optimizer state the first one allocates on a 24 GB card. The LIBERO page gains the invocation E11 evaluated with, copied from the harness script the run committed, and the flags that are easy to get wrong: --runner.pass-proc-id gives process i task i and makes --num-episodes a per-task count, --recorder.no-thread0-only is what makes per-task success rates recoverable afterwards, and rendering goes through EGL on a device EGL orders itself. The pi0 examples in four policy pages could not have run. They paired a flow policy with dppo, passed hopper - one of dppo's own presets - as that algorithm's config, and spelled every flag with underscores where the CLI generates hyphens. They are replaced by the eval and fpo invocations E11 ran. The underscore spellings elsewhere on the DPPO page are corrected as well: hyphenation is a property of the CLI, not of any one policy. Co-Authored-By: Claude Opus 5 (1M context) --- docs/env/libero_env.md | 35 +++++++++++++++++++++++++++++++ docs/env/libero_env.zh.md | 31 ++++++++++++++++++++++++++++ docs/index.md | 19 +++++++++++++++++ docs/index.zh.md | 15 ++++++++++++++ docs/policy/dppo_policy.md | 39 +++++++++++++++++++++++------------ docs/policy/dppo_policy.zh.md | 38 ++++++++++++++++++++++------------ docs/policy/index.md | 8 +++++-- docs/policy/index.zh.md | 8 +++++-- 8 files changed, 163 insertions(+), 30 deletions(-) diff --git a/docs/env/libero_env.md b/docs/env/libero_env.md index 9088043..285dbc2 100644 --- a/docs/env/libero_env.md +++ b/docs/env/libero_env.md @@ -24,6 +24,41 @@ Run a few episodes. plugrl-run-env-client libero-v1 --num-episodes 10 --server-host 127.0.0.1 --server-port 8000 ``` +## As E11 ran it + +E11 evaluated a `pi05_libero` checkpoint on LIBERO through a PlugRL server: one +client process per task, ten tasks at once, with each task's initial states +taken in order rather than sampled. + +```bash +plugrl-run-env-client libero-v1 \ + --server-host 127.0.0.1 --server-port 8000 \ + --num-envs 1 --num-procs 10 --num-episodes 10 \ + --env.task-suite-name libero_spatial \ + --env.no-randomize-initial-state \ + --runner.pass-proc-id \ + --runner.replan-steps 5 \ + --runner.seed 7 \ + --recorder.no-thread0-only \ + --exp-name my_eval +``` + +- `--runner.pass-proc-id` gives process *i* task *i*, so ten processes cover a + ten-task suite and `--num-episodes` counts per task rather than in total. To + hold one task fixed instead - which is what the fine-tuning runs did - drop + that flag and pass `--env.task-id 8 --env.randomize-initial-state`. +- `--recorder.no-thread0-only` makes every process write its own + `summary.json`. Without it only process 0 reports, and per-task success + rates cannot be recovered afterwards. +- Rendering goes through EGL, so the client needs `MUJOCO_GL=egl` and + `PYOPENGL_PLATFORM=egl` in its environment. Which GPU it renders on follows + EGL's own device order, which need not agree with `CUDA_VISIBLE_DEVICES`. +- LIBERO writes its own config file, so point `LIBERO_CONFIG_PATH` somewhere + writable before the first run. + +Protocol, results and the recorded environment of both processes: +[E11](https://github.com/PlugRL/plugrl-server/tree/main/experiments/e11-vla-rl-libero). + ## Verify - Env client can create the Libero env. diff --git a/docs/env/libero_env.zh.md b/docs/env/libero_env.zh.md index 4346c0b..7185d78 100644 --- a/docs/env/libero_env.zh.md +++ b/docs/env/libero_env.zh.md @@ -24,6 +24,37 @@ plugrl-run-env-client libero-v1 --help plugrl-run-env-client libero-v1 --num-episodes 10 --server-host 127.0.0.1 --server-port 8000 ``` +## E11 是怎么跑的 + +E11 通过一个 PlugRL server 在 LIBERO 上评测 `pi05_libero` checkpoint:一个任务一个 +客户端进程,十个任务同时跑,每个任务的初始状态按顺序取,而不是随机采样。 + +```bash +plugrl-run-env-client libero-v1 \ + --server-host 127.0.0.1 --server-port 8000 \ + --num-envs 1 --num-procs 10 --num-episodes 10 \ + --env.task-suite-name libero_spatial \ + --env.no-randomize-initial-state \ + --runner.pass-proc-id \ + --runner.replan-steps 5 \ + --runner.seed 7 \ + --recorder.no-thread0-only \ + --exp-name my_eval +``` + +- `--runner.pass-proc-id` 让第 *i* 个进程跑第 *i* 个任务,十个进程正好覆盖十任务 + 套件;此时 `--num-episodes` 是**每个任务**的回合数,不是总数。若想固定跑单个 + 任务——微调那几次就是这么做的——去掉该 flag,改传 + `--env.task-id 8 --env.randomize-initial-state`。 +- `--recorder.no-thread0-only` 让每个进程都写出自己的 `summary.json`。不加它就只有 + 0 号进程上报,事后无法还原逐任务成功率。 +- 渲染走 EGL,客户端环境里需要 `MUJOCO_GL=egl` 与 `PYOPENGL_PLATFORM=egl`。它渲染 + 在哪块 GPU 上由 EGL 自己的设备顺序决定,不一定与 `CUDA_VISIBLE_DEVICES` 一致。 +- LIBERO 会写自己的配置文件,首次运行前把 `LIBERO_CONFIG_PATH` 指到可写目录。 + +协议、结果,以及两个进程各自被记录下来的环境: +[E11](https://github.com/PlugRL/plugrl-server/tree/main/experiments/e11-vla-rl-libero)。 + ## 验证 - env client 能创建 Libero 环境 diff --git a/docs/index.md b/docs/index.md index e5f0c83..6d9eacb 100644 --- a/docs/index.md +++ b/docs/index.md @@ -57,6 +57,25 @@ confirm the two sides talk to each other, not to train anything. and starts stepping episodes. - With `fpo`, the server prints a metrics table whose `rollout/reward` rises. +## A real VLA, end to end + +The same two processes carry a full-size pi0.5. The env client steps LIBERO, +the server answers with actions, and FPO trains on the feedback that comes +back. Nothing about the boundary changes; only the policy does. + +As a control, the unmodified checkpoint scored 99 of 100 on `libero_spatial` +and 185 of 200 on `libero_10`, against openpi's published 98.8 and 92.4 - and +the server's episode and step counts matched the clients' exactly, which is +what says the transport dropped nothing. + +**The reinforcement learning result is negative.** One FPO iteration on the +hardest task took its success rate from 26 of 50 to 0 of 50, and the run is +incomplete at one iteration of ten: a second learn step does not fit beside +the optimizer state the first one allocates on a 24 GB card. The predictions +were pre-registered, and one of them is falsified. The numbers, the recorded +environment of both processes, and what none of it supports: +[E11](https://github.com/PlugRL/plugrl-server/tree/main/experiments/e11-vla-rl-libero). + ## Components - `plugrl-server`: training server, runs algorithm, policy, checkpoints, tracking diff --git a/docs/index.zh.md b/docs/index.zh.md index 4ed2462..e3e84c3 100644 --- a/docs/index.zh.md +++ b/docs/index.zh.md @@ -51,6 +51,21 @@ dummy 算法的 `learn` 是一个 sleep,不会移动任何权重。它用来 - env client 打印 server 元信息(策略名、动作形状)并开始跑 episode - 用 `fpo` 时,server 的指标表里 `rollout/reward` 会上升 +## 真实 VLA,端到端 + +同样的两个进程也能承载一个完整尺寸的 pi0.5:环境端步进 LIBERO,服务端返回动作, +FPO 用回传的反馈训练。边界本身没有任何改动,变的只是策略。 + +作为对照,未经微调的 checkpoint 在 `libero_spatial` 上 99/100、在 `libero_10` 上 +185/200,与 openpi 公布的 98.8 与 92.4 一致;而且服务端记录的回合数与步数和客户端 +逐一相等——正是这一条说明传输层没有悄悄丢掉任何东西。 + +**强化学习的结果是负面的。** 在最难的那个任务上,一轮 FPO 把成功率从 26/50 打到 +0/50;而且这次训练并不完整,10 轮只跑完 1 轮——第二次 learn 装不下,它要和第一次 +分配的优化器状态挤在同一张 24 GB 卡上。预测是预注册的,其中一条被证伪。数字、两个 +进程各自被记录下来的环境,以及这些数据**不能**支持的结论,见 +[E11](https://github.com/PlugRL/plugrl-server/tree/main/experiments/e11-vla-rl-libero)。 + ## 组件 - `plugrl-server`:训练端,负责算法、策略、checkpoint、指标追踪 diff --git a/docs/policy/dppo_policy.md b/docs/policy/dppo_policy.md index 640506e..bd4c1ff 100644 --- a/docs/policy/dppo_policy.md +++ b/docs/policy/dppo_policy.md @@ -13,15 +13,28 @@ This page covers: DPPO policy. ```bash -plugrl-run-server dppo-policy default dppo hopper --exp_name my_dppo_exp +plugrl-run-server dppo-policy default dppo hopper --exp-name my_dppo_exp ``` -OpenPI PI0 policy (requires a checkpoint directory). +OpenPI PI0 policy (requires a checkpoint directory). It is a flow policy: it +pairs with `fpo` or `eval`, never with `dppo`. These are the two invocations +[E11](https://github.com/PlugRL/plugrl-server/tree/main/experiments/e11-vla-rl-libero) ran. ```bash -plugrl-run-server pi0-policy default dppo hopper \ - --policy.checkpoint_path /path/to/pi0_checkpoint \ - --policy.name pi05_tiny_libero +# evaluation - no learning +plugrl-run-server pi0-policy default eval default \ + --policy.name pi05_libero \ + --policy.checkpoint-path /path/to/pi0_checkpoint \ + --policy.device cuda + +# FPO fine-tuning +plugrl-run-server pi0-policy default fpo default \ + --policy.name pi05_libero \ + --policy.checkpoint-path /path/to/pi0_checkpoint \ + --policy.device cuda \ + --algo.learning-rate 1e-5 --algo.batch-size 8 \ + --algo.n-samples-per-action 4 --algo.buffer-size 4096 \ + --algo.global-steps 40960 ``` ## Verify @@ -52,26 +65,26 @@ python -c "import my_pkg.plugrl_policies; from plugrl_server.cli import main; ma Common flags. -- `--policy.env_type gym` -- `--policy.env_name hopper-medium-v2` -- `--policy.checkpoint_path /path/to/checkpoint.pt` +- `--policy.env-type gym` +- `--policy.env-name hopper-medium-v2` +- `--policy.checkpoint-path /path/to/checkpoint.pt` - `--policy.critic.*` ## Built-in: `pi0-policy` (OpenPI) - UID: `pi0-policy` - Code: `plugrl-server/src/plugrl_server/policy/openpi/openpi_policy.py` -- `--policy.checkpoint_path` is required and must point to a directory with: +- `--policy.checkpoint-path` is required and must point to a directory with: - `model.safetensors` - `assets/` with normalization stats - Setup notes: `plugrl-server/src/plugrl_server/policy/openpi/README.md`. Common flags. -- `--policy.name pi05_tiny_libero` -- `--policy.denoising_steps 5` -- `--policy.train_expert_only true` -- `--policy.default_prompt "..."` +- `--policy.name pi05_libero` +- `--policy.denoising-steps 5` +- `--policy.train-expert-only true` +- `--policy.default-prompt "..."` ## Implement a diffusion-style policy diff --git a/docs/policy/dppo_policy.zh.md b/docs/policy/dppo_policy.zh.md index 453f72e..9149d3d 100644 --- a/docs/policy/dppo_policy.zh.md +++ b/docs/policy/dppo_policy.zh.md @@ -13,15 +13,27 @@ DPPO 策略。 ```bash -plugrl-run-server dppo-policy default dppo hopper --exp_name my_dppo_exp +plugrl-run-server dppo-policy default dppo hopper --exp-name my_dppo_exp ``` -OpenPI PI0 策略(需要 checkpoint 目录)。 +OpenPI PI0 策略(需要 checkpoint 目录)。它是流策略:与 `fpo` 或 `eval` 搭配, +**绝不**与 `dppo` 搭配。下面两条就是 [E11](https://github.com/PlugRL/plugrl-server/tree/main/experiments/e11-vla-rl-libero) 实际跑的命令。 ```bash -plugrl-run-server pi0-policy default dppo hopper \ - --policy.checkpoint_path /path/to/pi0_checkpoint \ - --policy.name pi05_tiny_libero +# 评测 —— 不训练 +plugrl-run-server pi0-policy default eval default \ + --policy.name pi05_libero \ + --policy.checkpoint-path /path/to/pi0_checkpoint \ + --policy.device cuda + +# FPO 微调 +plugrl-run-server pi0-policy default fpo default \ + --policy.name pi05_libero \ + --policy.checkpoint-path /path/to/pi0_checkpoint \ + --policy.device cuda \ + --algo.learning-rate 1e-5 --algo.batch-size 8 \ + --algo.n-samples-per-action 4 --algo.buffer-size 4096 \ + --algo.global-steps 40960 ``` ## 验证 @@ -52,26 +64,26 @@ python -c "import my_pkg.plugrl_policies; from plugrl_server.cli import main; ma 常用参数。 -- `--policy.env_type gym` -- `--policy.env_name hopper-medium-v2` -- `--policy.checkpoint_path /path/to/checkpoint.pt` +- `--policy.env-type gym` +- `--policy.env-name hopper-medium-v2` +- `--policy.checkpoint-path /path/to/checkpoint.pt` - `--policy.critic.*` ## 内置:`pi0-policy`(OpenPI) - UID:`pi0-policy` - 代码:`plugrl-server/src/plugrl_server/policy/openpi/openpi_policy.py` -- `--policy.checkpoint_path` 必填,目录内需要: +- `--policy.checkpoint-path` 必填,目录内需要: - `model.safetensors` - `assets/`(归一化统计) - 本地安装/替换步骤见:`plugrl-server/src/plugrl_server/policy/openpi/README.md`。 常用参数。 -- `--policy.name pi05_tiny_libero` -- `--policy.denoising_steps 5` -- `--policy.train_expert_only true` -- `--policy.default_prompt "..."` +- `--policy.name pi05_libero` +- `--policy.denoising-steps 5` +- `--policy.train-expert-only true` +- `--policy.default-prompt "..."` ## 自定义 diffusion policy diff --git a/docs/policy/index.md b/docs/policy/index.md index 6f76055..c367b2e 100644 --- a/docs/policy/index.md +++ b/docs/policy/index.md @@ -24,10 +24,14 @@ plugrl-run-server --help - `fpo-policy`: FPO flow-matching policy, used by the get-started run - `pi0-policy`: OpenPI policy, requires a checkpoint path -OpenPI example. +OpenPI example. `pi0-policy` is a flow policy, so it pairs with `fpo` or with +`eval` - not with `dppo`, which expects a diffusion policy. ```bash -plugrl-run-server pi0-policy default dppo hopper --policy.checkpoint_path /path/to/checkpoint +plugrl-run-server pi0-policy default eval default \ + --policy.name pi05_libero \ + --policy.checkpoint-path /path/to/checkpoint \ + --policy.device cuda ``` ## Troubleshooting diff --git a/docs/policy/index.zh.md b/docs/policy/index.zh.md index a4bf2a1..88239f5 100644 --- a/docs/policy/index.zh.md +++ b/docs/policy/index.zh.md @@ -24,10 +24,14 @@ plugrl-run-server --help - `fpo-policy`:FPO flow matching 策略,快速上手那条命令用的就是它 - `pi0-policy`:OpenPI 策略,需要 checkpoint 路径 -OpenPI 示例。 +OpenPI 示例。`pi0-policy` 是流策略,因此与 `fpo` 或 `eval` 搭配, +**不能**与 `dppo` 搭配——后者要的是 diffusion 策略。 ```bash -plugrl-run-server pi0-policy default dppo hopper --policy.checkpoint_path /path/to/checkpoint +plugrl-run-server pi0-policy default eval default \ + --policy.name pi05_libero \ + --policy.checkpoint-path /path/to/checkpoint \ + --policy.device cuda ``` ## 常见问题