I build multimodal systems that see, reason, retrieve, and act.
🎵 Play a tiny welcome chime · click to listen
- Undergraduate in the College of Electrical Engineering, Zhejiang University
- Exploring reliable multimodal intelligence, from visual tokens to tool-using agents
- Based in Hangzhou, China
| Vision-Language Models Visual token compression, long-video understanding, and grounded evaluation. |
Agentic Systems Planning, retrieval, tool use, and verification loops that catch mistakes. |
| World Models Learning dynamics that make imagined states useful for planning. |
Generative Models Controllable diffusion, autoregressive generation, and synthetic data. |
| Project | What it is |
|---|---|
| 🎮 game-ai-benchmarks-papers | Bilingual papers and benchmarks for AI game generation and interactive worlds. |
| 🧩 awesome-multi-reference-agentic-image-generation | Source-checked reading list on multi-reference and agentic image generation. |
| 🎴 algorithm-cards | 54 visual algorithm cards with diagrams, C++/Python templates, and two slide decks. |
| 🎓 zju-beamer-skill | An AI-invokable Zhejiang University Beamer workflow for polished academic slides. |
| 🧠 textbook-integration-agent | Knowledge-graph and RAG agent for integrating textbooks and answering questions. |
Open to research conversations and thoughtful collaborations · 3240102222@zju.edu.cn · zsyverse.github.io
