The challenge is structured in two rounds. In Phase 1, all teams are evaluated on a standardized real-robot bimanual manipulation benchmark using a shared dataset. In Phase 2, the top teams are selected and supported in iterative rollout collection to further improve their policies via post-training.
Phase 1 — Evaluation TracksPhase 1 evaluates submitted policies on 3–5 standardized real-robot tasks across two ranking tracks. Teams perform offline training on the released dataset (expert data + baseline success and failure rollouts) and submit policies for benchmark evaluation.
Policies are evaluated and ranked on each task independently. Best for specialists optimizing for individual task performance.
A single policy is evaluated jointly across all tasks. Best for generalist policies that share representations across skills.
Phase 1 includes three real-robot bimanual manipulation tasks. Select a task below to view a sample rollout video and its scoring criteria.
insert-mouse-batterytower-of-hanoi-gameseal-water-bottle-capFor Phase 1, we release a standardized real-robot bimanual dataset designed for offline training. The dataset contains four complementary components:
High-quality human teleoperation demonstrations on the benchmark tasks.
Trajectories where a baseline policy successfully completed the task.
Trajectories where the baseline policy failed — useful negative signal for post-training.
Human interventions, corrections, and preference labels collected during baseline rollouts.
All four components are intended for offline training in Phase 1. Hosted on Hugging Face Datasets.
Download on Hugging FaceStep-by-step instructions for joining the challenge — including environment setup, dataset access, baseline reproduction, evaluation protocol, and submission format — are hosted in our GitHub repositories. Reference code and starter scripts are provided so teams can get up and running quickly.
Based on Phase 1 results, the top 3 teams advance to Phase 2. During this round, selected teams collect rollouts by deploying their own policies and use those rollouts to further improve performance via post-training.
Listed prizes apply to each track (Single-Task & Multi-Task).
Final Phase 1 evaluation of all submitted policies on the three benchmark tasks. The highlighted top teams advanced to Phase 2.
| # | Team / Model | insert-mouse-battery |
tower-of-hanoi-game |
seal-water-bottle-cap |
Average | ||||
|---|---|---|---|---|---|---|---|---|---|
| Score | SR | Score | SR | Score | SR | Score | SR | ||
| 1 |
Haruki
U-Tokyo
trials: 10 / 10 / 10
|
90 | 90% | 66 | 50% | 90 | 90% | 82 | 76.7% |
| 2 |
PengfangQian
Fudan & SII
trials: 10 / 10 / 10
|
90 | 90% | 33 | 30% | 80 | 60% | 67.7 | 60% |
| 3 |
VLAlab-JP
trials: 10 / 10 / 10
|
100 | 100% | 40 | 40% | 30 | 30% | 56.7 | 56.7% |
| 4 |
Zhangyu
Tsinghua
trials: 10 / 10 / 10
|
86 | 70% | 59 | 50% | 38 | 30% | 61 | 50% |
| 5 |
TongJiang
cityu
trials: 4 / 5 / 5
|
70 | 50% | 25 | 20% | 70 | 60% | 53.9 | 42.9% |
| 6 |
JiabingYang
CASIA
trials: 5 / 5 / 5
|
100 | 100% | 12 | 0% | 40 | 20% | 50.7 | 40% |
| 7 |
Benson
Tongji
trials: 5 / 5 / 5
|
60 | 60% | 6 | 0% | 60 | 60% | 42 | 40% |
| 8 |
RuikaiShi
HKUST
trials: 5 / 5 / 5
|
80 | 80% | 0 | 0% | 40 | 40% | 40 | 40% |
| 9 |
HoKyunIm
Yonsei
trials: 5 / 5 / 5
|
60 | 60% | 0 | 0% | 40 | 40% | 33.3 | 33.3% |
| 10 |
Yitong
SII
trials: 5 / 10 / 10
|
80 | 80% | 9 | 0% | 46 | 40% | 38 | 32% |
| 11 |
ShijieGeng
Drexel
trials: 5 / 5 / 5
|
52 | 20% | 7.5 | 0% | 60 | 60% | 39.8 | 26.7% |
| 12 |
Guanqi
HKU
trials: 10 / 10 / 5
|
66 | 50% | 6 | 0% | 52 | 20% | 41.3 | 23.3% |
| 13 |
QianYe
jingshuo
trials: 10 / 10 / 5
|
56 | 20% | 6 | 0% | 40 | 40% | 32.8 | 16% |
| 14 |
Jingyu
SYSU
trials: 10 / 10 / 10
|
72 | 40% | 7.8 | 0% | 14 | 0% | 32.1 | 13.3% |
| 15 |
Xingxin
HKUST
trials: 5 / 5 / 5
|
40 | 40% | 18 | 0% | 0 | 0% | 19.3 | 13.3% |
| 16 |
Yangxin
SYSU
trials: 5 / 5 / 5
|
40 | 40% | 0 | 0% | 0 | 0% | 13.3 | 13.3% |
| 17 |
Qichang
SYSU
trials: 5 / 5 / 2
|
16 | 0% | 0 | 0% | 0 | 0% | 5.3 | 0% |
| 18 |
SiyangZheng
SYSU
trials: 5 / — / —
|
8 | 0% | — | — | — | — | 2.7 | 0% |
| 19 |
Zhihaozhan
SYSU
trials: 5 / 5 / 5
|
0 | 0% | 0 | 0% | 0 | 0% | 0 | 0% |
| # | Team / Model | insert-mouse-battery |
tower-of-hanoi-game |
seal-water-bottle-cap |
Average | ||||
|---|---|---|---|---|---|---|---|---|---|
| Score | SR | Score | SR | Score | SR | Score | SR | ||
| 1 |
Haruki
U-Tokyo
trials: 10 / 10 / 10
|
80 | 80% | 66 | 60% | 60 | 60% | 68.7 | 66.7% |
| 2 |
Zecheng
ICL & Psibot
trials: 10 / 10 / 10
|
80 | 80% | 52 | 40% | 55 | 50% | 62.3 | 56.7% |
| 3 |
VLAlab-JP
trials: 13 / 13 / 13
|
98.5 | 92.3% | 48.5 | 46.2% | 34.6 | 23.1% | 60.5 | 53.8% |
| 4 |
PengfangQian
Fudan & SII
trials: 10 / 12 / 10
|
96 | 80% | 44.2 | 41.7% | 40 | 40% | 59.1 | 53.1% |
| 5 |
TongJiang
cityu
trials: 12 / 13 / 13
|
71.7 | 58.3% | 23.1 | 23.1% | 69.2 | 61.5% | 54.2 | 47.4% |
| 6 |
Zhangyu
Tsinghua
trials: 10 / 10 / 10
|
98 | 90% | 35 | 20% | 35 | 30% | 56 | 46.7% |
| 7 |
HanZhao
Westlake
trials: 10 / 5 / 5
|
64 | 60% | 20 | 20% | 70 | 60% | 51.3 | 46.7% |
| 8 |
ZeyuPing
SYSU
trials: 5 / 5 / 5
|
60 | 60% | 46 | 40% | 30 | 20% | 45.3 | 40% |
| 9 |
Xingxin
HKUST
trials: 10 / 10 / 10
|
50 | 50% | 14.4 | 10% | 60 | 50% | 41.5 | 36.7% |
| 10 |
ShijieGeng
Drexel
trials: 5 / 5 / 5
|
76 | 60% | 0 | 0% | 54 | 40% | 43.3 | 33.3% |
| 11 |
JiabingYang
CASIA
trials: 5 / 5 / 5
|
80 | 80% | 0 | 0% | 20 | 20% | 33.3 | 33.3% |
| 12 |
Guanqi
HKU
trials: 4 / 10 / 5
|
70 | 50% | 38.9 | 20% | 32 | 20% | 47 | 30% |
| 13 |
HoKyunIm
Yonsei
trials: 10 / 10 / 10
|
68 | 60% | 10 | 10% | 27 | 20% | 35 | 30% |
| 14 |
Yitong
SII
trials: 10 / 10 / 10
|
40 | 20% | 10 | 10% | 71.1 | 50% | 40.4 | 26.7% |
| 15 |
Jingyu
SYSU
trials: 5 / 5 / 5
|
80 | 80% | 0 | 0% | 0 | 0% | 26.7 | 26.7% |
| 16 |
HaotongChen
Tongji
trials: 3 / 3 / 3
|
33.3 | 33.3% | 33.3 | 33.3% | 0 | 0% | 22.2 | 22.2% |
| 17 |
QianYe
jingshuo
trials: 5 / 5 / 10
|
80 | 80% | 12 | 0% | 5 | 0% | 25.5 | 20% |
| 18 |
Zhihaozhan
SYSU
trials: 5 / 5 / —
|
16 | 0% | 38 | 20% | — | — | 18 | 6.7% |
| 19 |
Qichang
SYSU
trials: 5 / 5 / 2
|
24 | 0% | 7.5 | 0% | 0 | 0% | 10.5 | 0% |
| 20 |
TianxingShi
Tongji
trials: 5 / 5 / 5
|
8 | 0% | 0 | 0% | 0 | 0% | 2.7 | 0% |
Score: progress score per task (higher is better). SR: success rate over rollouts. trials: number of evaluation rollouts per task, listed in column order insert-mouse-battery / tower-of-hanoi-game / seal-water-bottle-cap. Ranking is by Average SR first, then Average Score as a tie-breaker. In both tracks the Average is computed over all three tasks — any task a team did not evaluate (—) counts as 0. Runs with only a single rollout are excluded.
Phase 1 final results. These are the final Phase 1 evaluation results. Each policy was evaluated with up to 5 rollouts per task under a unified setting.
Phase 2 teams deploy their own policies to collect rollouts and improve them iteratively via post-training. Every iteration (iter 1–3) is evaluated on the same three tasks and reported below, with each team's Phase 1 result shown as the baseline. The table is updated as new iterations are evaluated.
| # | Team / Model | Run | insert-mouse-battery |
tower-of-hanoi-game |
seal-water-bottle-cap |
Average | ||||
|---|---|---|---|---|---|---|---|---|---|---|
| Score | SR | Score | SR | Score | SR | Score | SR | |||
| 🥇1 |
Haruki
U-Tokyo
trials: 10 / 10 / 10
|
Phase 1 | 90 | 90% | 66 | 50% | 90 | 90% | 82 | 76.7% |
| 2 |
Haruki
U-Tokyo
trials: 15 / 15 / 12
|
iter 2 | 100 | 100% | 63 | 50% | 60 | 60% | 74.3 | 70% |
| 🥈3 |
VLAlab-JP
trials: 15 / 11 / 15
|
iter 1 | 86 | 70% | 63 | 60% | 57 | 50% | 68.7 | 60% |
| 🥉4 |
PengfangQian
Fudan & SII
trials: 10 / 10 / 10
|
Phase 1 | 90 | 90% | 33 | 30% | 80 | 60% | 67.7 | 60% |
| 5 |
VLAlab-JP
trials: 10 / 10 / 10
|
Phase 1 | 100 | 100% | 40 | 40% | 30 | 30% | 56.7 | 56.7% |
| 6 |
Zhangyu
Tsinghua
trials: 10 / 10 / 10
|
Phase 1 | 86 | 70% | 59 | 50% | 38 | 30% | 61 | 50% |
| 7 |
Haruki
U-Tokyo
trials: 15 / 15 / 15
|
iter 1 | 74 | 70% | 63 | 60% | 39 | 20% | 58.7 | 50% |
| 8 |
Zhangyu
Tsinghua
trials: 15 / 15 / 15
|
iter 1 | 100 | 100% | 36 | 30% | 31 | 20% | 55.7 | 50% |
| 9 |
PengfangQian
Fudan & SII
trials: 15 / 15 / 10
|
iter 2 | 16 | 0% | 56 | 50% | 53 | 50% | 41.7 | 33.3% |
| 10 |
VLAlab-JP
trials: 12 / 12 / 12
|
iter 3 | 68 | 40% | 0 | 0% | 55 | 50% | 41 | 30% |
| 11 |
VLAlab-JP
trials: 15 / 10 / 15
|
iter 2 | 40 | 20% | 9 | 0% | 72 | 60% | 40.3 | 26.7% |
| Not fully evaluated — tasks with fewer than 10 rollouts count as 0 in the average | ||||||||||
| — |
PengfangQian
Fudan & SII
trials: 15 / 15 / —
|
iter 3 | 94 | 90% | 58 | 40% | — | — | 50.7 | 43.3% |
| — |
Zhangyu
Tsinghua
trials: 10 / 10 / —
|
iter 2 | 92 | 80% | 3 | 0% | — | — | 31.7 | 26.7% |
| — |
Haruki
U-Tokyo
trials: — / 12 / 12
|
iter 3 | — | — | 30 | 10% | 74 | 60% | 34.7 | 23.3% |
| — |
PengfangQian
Fudan & SII
trials: 11 / 14 / 4
|
iter 1 | 2 | 0% | 3 | 0% | 25 | 25% | 1.7 | 0% |
| — |
Zhangyu
Tsinghua
trials: — / 10 / —
|
iter 3 | — | — | 0 | 0% | — | — | 0 | 0% |
| # | Team / Model | Run | insert-mouse-battery |
tower-of-hanoi-game |
seal-water-bottle-cap |
Average | ||||
|---|---|---|---|---|---|---|---|---|---|---|
| Score | SR | Score | SR | Score | SR | Score | SR | |||
| 🥇1 |
Haruki
U-Tokyo
trials: 10 / 10 / 10
|
Phase 1 | 80 | 80% | 66 | 60% | 60 | 60% | 68.7 | 66.7% |
| 🥈2 |
PengfangQian
Fudan & SII
trials: 15 / 15 / 13
|
iter 1 | 100 | 100% | 72 | 50% | 37 | 30% | 69.7 | 60% |
| 3 |
Haruki
U-Tokyo
trials: 15 / 15 / 12
|
iter 3 | 70 | 70% | 90 | 90% | 42 | 20% | 67.3 | 60% |
| 🥉4 |
VLAlab-JP
trials: 15 / 13 / 15
|
iter 1 | 60 | 60% | 75 | 60% | 66 | 60% | 67 | 60% |
| 5 |
Haruki
U-Tokyo
trials: 15 / 15 / 15
|
iter 1 | 88 | 80% | 83 | 80% | 20 | 20% | 63.7 | 60% |
| 6 |
Zecheng
ICL & Psibot
trials: 10 / 10 / 10
|
Phase 1 | 80 | 80% | 52 | 40% | 55 | 50% | 62.3 | 56.7% |
| 7 |
VLAlab-JP
trials: 13 / 13 / 13
|
Phase 1 | 98.5 | 92.3% | 48.5 | 46.2% | 34.6 | 23.1% | 60.5 | 53.8% |
| 8 |
PengfangQian
Fudan & SII
trials: 10 / 12 / 10
|
Phase 1 | 96 | 80% | 44.2 | 41.7% | 40 | 40% | 59.1 | 53.1% |
| 9 |
Haruki
U-Tokyo
trials: 15 / 15 / 15
|
iter 2 | 80 | 80% | 43 | 40% | 31 | 20% | 51.3 | 46.7% |
| 10 |
VLAlab-JP
trials: 15 / 14 / 15
|
iter 2 | 52 | 40% | 25 | 10% | 85 | 80% | 54 | 43.3% |
| 11 |
Zecheng
ICL & Psibot
trials: 15 / 15 / 15
|
iter 1 | 70 | 70% | 38 | 20% | 46 | 20% | 51.3 | 36.7% |
| 12 |
VLAlab-JP
trials: 15 / 11 / 12
|
iter 3 | 54 | 50% | 10 | 10% | 20 | 20% | 28 | 26.7% |
| Not fully evaluated — tasks with fewer than 10 rollouts count as 0 in the average | ||||||||||
| — |
PengfangQian
Fudan & SII
trials: 15 / 15 / —
|
iter 2 | 88 | 80% | 60 | 50% | — | — | 49.3 | 43.3% |
Each row is one evaluated run: a team's Phase 1 baseline or a Phase 2 iteration (iter 1–3). Score / SR are computed over the first 10 evaluation rollouts per task; rollouts beyond 10 serve only as a tie-breaker. Ranking is by Average SR, then Average Score. Medals (🥇🥈🥉) mark the top-3 teams, placed on each team's best run. Runs below the divider were not fully evaluated — tasks with fewer than 10 rollouts (or not evaluated, —) count as 0 in the average.
Team registration is now open. Please fill in the Google Form below with your team information and track preference. Submission instructions will be shared with registered teams.
Register your team
Abstract. Traditionally, policy training pipelines are split into two phases: a pre-training phase, that instills generalization, and a post-training phase, that tunes for maximum performance. The best post-trained policies are often created through skillful data curation, which discards a large fraction of the available training data and leaves a lot of robustness on the table. In this talk, I will discuss our recent work on achieving post-training performance without post-training in the pi07 model. Through careful conditioning of the pre-training process, we can train policies that match the performance of our best post-training pipelines, while retaining the full generality of the pre-training process.
Bio. Karl Pertsch is a member of the technical staff at Physical Intelligence. Before, he was a postdoc at UC Berkeley and Stanford, and obtained his PhD from USC. His work focuses on building generalist robot policies that can solve a wide range of physical manipulation tasks in the real world. His work has been awarded the Best Conference Paper Award at ICRA'24, two Outstanding Paper Awards Finalists at CoRL'24, and a Best Paper Finalist at RSS'25.
Abstract. Robotic post-training via reinforcement learning, while promising, has yet to see the same success that has been observed in domains like large language modeling. There is a considerable difference in assumptions in post-training between the robotics and language domains, making many standard RL algorithms ineffective. We posit that using simulation for post-training in a real-to-sim-to-real loop can bridge this gap, allowing for cheap, fast data collection for post-training. But the use of simulation for post-training must contend with the inevitable gap between imperfect simulation and reality, both from real-to-sim and from sim-to-real. We will present a simple way to bridge this gap, enabling post-training via RL to significantly improve the success rate, throughput and robustness of robotic foundation models. We will end with broader perspectives on the role of simulation in post-training of generalist robotic policies.
Bio. Abhishek Gupta is an assistant professor in the Paul G. Allen School of Computer Science and Engineering at the University of Washington since 2022. He leads the Washington Embodied Intelligence and Robotics Development lab focusing on robot learning and reinforcement learning. Previously, he was a postdoctoral scholar at MIT, collaborating with Russ Tedrake and Pulkit Agrawal. Prior to that he received his Ph.D. and B.S. degrees from UC Berkeley, working with Sergey Levine and Pieter Abbeel. Abhishek is the recipient of the IEEE RAS Early Career Award, the Toyota Research Institute Young Investigator award and an Amazon Science Hub award, along with award nominations at several top conferences and workshops. His research interests lie in scalable reinforcement learning methods for robot learning, in particular methods for continual adaptation in the real world.
Abstract. Robotic foundation models are emerging as a scalable path toward generalist robots. Large-scale pretraining provides the foundation for broad representations and initial capabilities, while post-training on real-world deployment experience refines and expands these capabilities toward reliable operation. In this talk, I will present a closed-loop approach to physical AI through two complementary efforts. First, I will introduce tau-0, an open-source robot foundation model trained on diverse robot and non-robot data to learn predictive representations of physical interactions. Second, I will discuss Learning While Deploying, a framework that turns deployment into a continual post-training process, converting real-world successes and failures into policy improvement through real-world learning. Together, these components form a data flywheel for physical AI: pretrain from large-scale diverse data, deploy in real environments, learn from interaction, and continuously improve robot capabilities through real-world experience.
Bio. Jianlan Luo is an Associate Professor at the Shanghai Innovation Institute and Chief Scientist at AGIBOT. He received his Ph.D. from the University of California, Berkeley. After completing his Ph.D., he worked as a researcher at Google before returning to UC Berkeley as a postdoctoral scholar. His research focuses on building principled and scalable robotic learning systems that enable reliable, high-performance behavior in complex real-world environments. His work has been recognized by honors including MIT Technology Review's TR35 China, and has been featured in media outlets including WIRED, TechCrunch and more.
July 13, 2026 · 8:30 – 12:30 (Sydney time, AEST)
| Time | Event |
|---|---|
| 8:30 - 8:40 | Opening Remarks |
| 8:40 - 9:10 | Invited Talk 1 · Karl Pertsch |
| 9:10 - 9:40 | Invited Talk 2 · Abhishek Gupta |
| 9:40 - 10:10 | Invited Talk 3 · Jianlan Luo |
| 10:10 - 10:40 | Coffee Chat & Sponsor Showcase |
| 10:40 - 11:40 | Participating Team Presentations |
| 11:40 - 12:10 | Panel Discussion & Audience Q&A |
| 12:10 - 12:30 | Closing & Summary |
Discussion topics for this workshop include, but are not limited to:
If you use this dataset or reference the challenge in your research, please cite us:
@misc{posttraining_robotics_2026,
title = {Post-Training for Robotics Foundation Models Dataset and Challenge},
author = {Zhang, Shiduo and Wang, Yue and Chang, Haonan and Zhao, Hang and
Liu, Yicheng and Guizilini, Vitor and Bobu, Andreea and Wagenmaker, Andrew and
Dixit, Anushri and Yu, Chao and Shah, Dhruv and Simchowitz, Max},
year = {2026},
howpublished = {RSS 2026 Workshop & Challenge},
url = {https://posttraining-for-robotics.github.io}
}
For any questions about the workshop or the challenge, please reach out to:
Shiduo Zhang · sdzhang23@m.fudan.edu.cn