From 4eb3f534057be31ef9c27f68e2b69f5654dcb17e Mon Sep 17 00:00:00 2001 From: tactino <18781106300@163.com> Date: Fri, 18 Sep 2026 16:48:02 -0400 Subject: [PATCH] docs: put the robot above the numbers A clip of the unmodified pi0.5 on libero_spatial task 0, three episodes, all successful, captured through this server with the env client's recorder on. The frames are the observation stream, not an outside camera. It is served from the documentation site rather than committed here: the clip is 934 KB at the policy's native 224x224 and this repository caps added files at 500 KB, so carrying it would mean a choppier or smaller loop for no reason. That makes this depend on plugrl.github.io#9 - merged before that deploys, the image is broken. The caption names it as the baseline; the fine-tuned policy is the one that scores zero further down the page. --- README.md | 8 ++++++++ 1 file changed, 8 insertions(+) diff --git a/README.md b/README.md index 6842a48..e0636bb 100644 --- a/README.md +++ b/README.md @@ -7,6 +7,14 @@ The training side of PlugRL: it holds the policy and the learning algorithm, batches inference across every connected env client over WebSocket, and learns from the feedback those clients send back. +A Franka arm in LIBERO reaching for and picking up an object, seen from the policy's own camera + +*The **unmodified** pi0.5, driven through this server on `libero_spatial` task +0 - three episodes, all three successful. These are the frames the env client +sends as observations, at the policy's native 224x224, not an outside camera. +The fine-tuned policy is the one that scores zero, below.* + **A full-size pi0.5 has run end to end through it on LIBERO** - inference, feedback and FPO training - with the server's record of episodes and steps reconciling exactly with the clients'. As a control, the unmodified checkpoint