Embodied Decision Intelligence Lab (EDI Lab) 清华大学具身决策智能实验室

Introduction

My name is Chao Yu (于超). I received my Ph.D. from the Department of Electronic Engineering at Tsinghua University in 2023. I am currently an Assistant Professor (Distinguished Research Fellow) at the Embodied Decision Intelligence Lab (EDI Lab) at SIGS of Tsinghua University. The EDI Lab is a sub-group of the Nanoscale Integrated Circuits and System Lab, Energy Efficient Computing Group (NICS-EFC) at Tsinghua University’s main campus. I also serve as the chairman of the SIGS of Tsinghua University - AgiBot Joint Research Center for Embodied Cognition and Decision Systems (JCES) and the co-founder of Striding AI (正行创新). I have been selected for the Youth Talent Support Program of the Chinese Institute of Electronics. My research has long focused on reinforcement learning-based decision intelligence. As first author or corresponding author, I have published more than 50 papers in top-tier international conferences and journals, including ICML, NeurIPS, ICLR, CVPR, ECCV, CoRL, IROS, ICRA, TMLR, and RAL, with over 7,000 citations on Google Scholar. My representative works include the multi-agent reinforcement learning algorithm MAPPO, which has received more than 4,000 Google Scholar citations, and RLinf, a large-scale reinforcement learning training framework for embodied intelligence, which has accumulated over 4,000 GitHub stars.

Feel free to reach out if you’d like to discuss research or explore potential collaboration!

Our lab is currently recruiting Master’s students, Ph.D. students, joint-program Ph.D. students with Zhongguancun Academy, postdoctoral researchers, and undergraduate research assistants. We warmly welcome students who are interested in recommended admission or applying through the graduate entrance examination to SIGS programs such as Artificial Intelligence, Data Science and Information Technology, Big Data Engineering, and Electronic Engineering, as well as applicants for the above Ph.D. and postdoctoral positions, to join us.

Highlights

VS-Bench: Evaluating VLMs for Strategic Reasoning and Decision-Making in Multi-Agent Environments
VS-Bench: Evaluating VLMs for Strategic Reasoning and Decision-Making in Multi-Agent Environments
Zelai Xu, Zhexuan Xu, Xiangmin Yi, Huining Yuan, Xinlei Chen, Yi Wu, Chao Yu, Yu Wang
IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR 2026), Oral  ·  2026
RLinf: Flexible and Efficient Large-scale Reinforcement Learning via Macro-to-Micro Flow Transformation
RLinf: Flexible and Efficient Large-scale Reinforcement Learning via Macro-to-Micro Flow Transformation
Chao Yu, Yuanqing Wang, Zhen Guo, Hao Lin, Si Xu, Hongzhi Zang, Quanlu Zhang, Yongji Wu, Chunyang Zhu, Junhao Hu, Zixiao Huang, Mingjie Wei, Yuqing Xie, Ke Yang, Bo Dai, Zhexuan Xu, Xiangyuan Wang, Xu Fu, Zhihao Liu, Kang Chen, Weilin Liu, Gang Liu, Boxun Li, Jianlei Yang, Zhi Yang, Guohao Dai, Yu Wang*
20th USENIX Symposium on Operating Systems Design and Implementation (OSDI 2026)  ·  2026
RLinf-VLA: A Unified and Efficient Framework for Reinforcement Learning of Vision-Language-Action Models
RLinf-VLA: A Unified and Efficient Framework for Reinforcement Learning of Vision-Language-Action Models
Hongzhi Zang, Mingjie Wei, Si Xu, Yongji Wu, Zhen Guo, Yuanqing Wang, Hao Lin, Liangzhi Shi, Yuqing Xie, Zhexuan Xu, Zhihao Liu, Kang Chen, Wenhao Tang, Quanlu Zhang, Weinan Zhang, Chao Yu, Yu Wang
Robotics: Science and Systems (RSS 2026)  ·  2026
Language Agents with Reinforcement Learning for Strategic Play in the Werewolf Game
Language Agents with Reinforcement Learning for Strategic Play in the Werewolf Game
Zelai Xu, Chao Yu, Fei Fang, Yu Wang, Yi Wu
International Conference on Machine Learning (ICML 2024)  ·  2024
Is DPO Superior to PPO for LLM Alignment? A Comprehensive Study
Is DPO Superior to PPO for LLM Alignment? A Comprehensive Study
Shusheng Xu, Wei Fu, Jiaxuan Gao, Wenjie Ye, Weilin Liu, Zhiyu Mei, Guangju Wang, Chao Yu, Yi Wu
International Conference on Machine Learning (ICML 2024), Oral  ·  2024
The Surprising Effectiveness of PPO in Cooperative Multi-Agent Games
The Surprising Effectiveness of PPO in Cooperative Multi-Agent Games
Chao Yu, Akash Velu, Eugene Vinitsky, Jiaxuan Gao, Yu Wang, Alexandre Bayen, Yi Wu
Conference on Neural Information Processing Systems (NeurIPS 2022), Datasets and Benchmarks Track  ·  2022

News

昆仑芯完成RLinf适配验证,拓展具身智能强化学习开源生态

RLinf 是面向具身智能的“仿训推一体化”大规模强化学习框架,由清华大学、无问芯穹联合发起,于 2025 年 9 月开源。框架面向具身智能与智能体场景,提供覆盖仿真、训练和推理的统一强化学习基础设施,打通渲染、训练、推理等环节,为开发者提供从仿真环境交互到策略更新的完整强化学习流程,目前已在机器人等相关场景中开展应用。

壁仞科技算力接入RLinf

得益于 RLinf 对硬件平台的抽象解耦设计,BIRENSUPA™ 的接入遵循标准化插件模式,对其他硬件平台代码零侵入。

清华大学联合无问芯穹开源具身智能云原生纳管平台RLark,5分钟接入、10秒启动任务

清华大学与无问芯穹共同打造并开源面向具身智能的云原生纳管平台 RLark。平台以自研的具身设备运行时(embodied-runtime)和任务级跨集群网络互联技术为核心,打通“真实设备如何被统一调度”与“跨地域任务如何互联运行”两个关键环节。

再也不用担心机器人硬件更新!用 RLinf 像搭乐高一样组装机器人

在真机强化学习实验中,更换夹爪、增设相机或扩展双臂平台通常需要修改多处代码。早期 RLinf 环境中的设备管理逻辑与特定硬件平台耦合,即使底层驱动可以复用,更换部件时仍需调整环境或控制器。

GPT-6 Astra 走向机器人,开源基础设施 RPent 正式上线了!

Codex、Claude Code 等数字智能体已经展示了一种高效的工作方式:大模型理解目标、拆解任务、调用工具,并根据结果持续调整。随着 GPT-6 Astra 开始接受真实机器人操作测试,同一种智能体范式正在进入物理世界。

Outreach

河套学院 AI Seminar The 55th Session: Embodied Intelligence Reinforcement Learning Training and Intelligent Infrastructure Exploration

于超老师受邀在深圳河套学院 AI Seminar 第 55 期作“Embodied Intelligence Reinforcement Learning Training and Intelligent Infrastructure Exploration”报告。

摩尔线程 MUSA 开源技术沙龙 RLinf × MUSA Meetup

于超老师受邀在 MUSA 开源技术沙龙 RLinf × MUSA Meetup 作“面向具身智能的高灵活大规模强化学习框架 RLinf”报告。

NVIDIA(英伟达)Invited Talk

于超老师受邀参加 NVIDIA(英伟达)Invited Talk。

RLinf 开源一周年研发回顾 Meetup

活动回顾

2026 GEIA 粤港澳大湾区具身智能与人形机器人创新周(深圳)

活动详情

Sponsors