Skip to content
View vigorlee's full-sized avatar
🎯
Focusing
🎯
Focusing

Highlights

  • Pro

Block or report vigorlee

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Content in all repositories owned by your account will be closed.
Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
vigorlee/README.md

Mingyi Li

PhD candidate in Vehicle Engineering at Beijing Institute of Technology (expected June 2027); research in Computer Science and Technology focused on robotic world models and generative policies

Email · LinkedIn · Repositories

I work on robotic world models and generative action policies, video pretraining and post-training, vision-language-action (VLA) systems, and robot navigation and manipulation. Recent work includes egocentric-video and robot-trajectory alignment on Unitree G1, staged egocentric-video pretraining for BAAI's EgoWorld-1, and VLA/GR00T-related mobile manipulation at LimX Dynamics. I have also led a small robotics solutions team and care about translating research into reliable applications. Reinforcement learning is not my primary specialization.

Research interests

  • Robotic world models and generative policies: predictive representations from video, simulation, and robot trajectories; action generation, long-horizon forecasting, and evaluation.
  • VLA and real-robot systems: connecting learned representations to navigation, manipulation, and physical-robot tests.
  • Multimodal memory: structured semantic-spatial retrieval for open-world object-goal navigation.
  • Reproducible systems: experimental infrastructure and transparent simulation versus real-robot evidence.

Selected publication

Selected projects

Multimodal knowledge graph integration for open-world object-goal navigation. This project investigates the use of structured memory and multimodal retrieval to support navigation decisions.

Research focus

Visual observations are organized into semantic-spatial memory that can be queried during navigation. The work concerns the connection between environmental observations, retrieved evidence, and object-goal decisions.

Source code and project documentation

World-model-assisted navigation and execution for the Go2-W platform, including map-independent charging and hybrid long-range navigation workflows.

System details and experimental materials

In the hybrid navigation workflow, Cosmos3-Edge selects an approved route, Nav2/RoamerX handles navigation, and DreamWaQ controls motion.

The September 10, 2026 simulation record covers stairs, ramps, and flat-ground obstacles with moving cylinder proxies. It reports 18 completed stages and 60.88 m of travel in MuJoCo / Unreal Engine, using sensor/model-fused obstacle data. These results describe the recorded simulation runs, rather than a real-robot benchmark.

Experimental records · Simulation overview

An experimental framework for persona agents with explicit memory and retrieval, with adversarial smoke tests for inspecting agent behavior.

Related infrastructure

Isaac Sim Kitchen provides a reproducible Lightwheel Kitchen scene setup. Room contains related runtime work. These repositories support scripted setup, asset checks, and experimental inspection.


Total repository stars

Current stars received across public repositories owned by this account. Automatically updated by the badge service; GitHub image caching may delay changes.

Popular repositories Loading

  1. wave-go wave-go Public

    WAVE-Go: Go2-W world-model action adaptation, verified mapless charging, and Cosmos3-Edge hybrid long-range navigation demos

    Python 105 1

  2. eidra-agent eidra-agent Public

    Persona agents with explicit memory, retrieval, adversarial smoke tests and a practical post-training roadmap.

    JavaScript 95

  3. Multimodal--RAG Multimodal--RAG Public

    EMKG: multimodal memory and retrieval for open-world object-goal navigation in dynamic environments.

    Python 93

  4. RoboKino RoboKino Public

    Isaac Sim-based dual-arm manipulation, expert data collection, and policy evaluation.

    Python 3

  5. SmarterStreaming SmarterStreaming Public

    Forked from daniulive/SmarterStreaming

    国内外为数不多致力于极致体验的超强全自研跨平台(windows/android/iOS)流媒体内核,通过模块化自由组合,支持实时RTMP推流、RTSP推流、RTMP播放器、RTSP播放器、录像、多路流媒体转发、音视频导播、动态视频合成、音频混音、直播互动、内置轻量级RTSP服务等,比快更快,业界真正靠谱的超低延迟直播SDK(1秒内,低延迟模式下200~400ms)。

    Java

  6. EasyPlayer-RTSP-Android EasyPlayer-RTSP-Android Public

    Forked from yiqideren/EasyPlayer_Android

    An elegant, simple, fast android RTSP/RTMP/HLS/HTTP Player.EasyPlayer support RTSP(RTP over TCP/UDP)version & Pro version,cover all kinds of streaming media!EasyPlayer是一款精炼、高效、稳定的流媒体播放器,分为RTSP版和Pro…

    Java