diff --git a/.gitignore b/.gitignore index 126722d..6345b91 100644 --- a/.gitignore +++ b/.gitignore @@ -23,6 +23,7 @@ wandb/ runs/ outputs/ checkpoints/ +/packages/ *.ckpt *.pt *.pth diff --git a/README.md b/README.md index 33a1e34..4aee340 100644 --- a/README.md +++ b/README.md @@ -119,14 +119,16 @@ setting. | Method | Clean ↑ | Random ↑ | Total ↑ | |---|---:|---:|---:| -| π₀ | 65.92 | 58.40 | 62.2 | -| π₀.₅ | 82.74 | 76.76 | 79.8 | -| Motus | **88.66** | 87.02 | **87.8** | -| Motus from WAN2.2 | 77.56 | 77.00 | 77.3 | -| FastWAM-Joint | 87.8 | 87.32 | 87.56 | -| StarWAM-Joint | 84.8 | 86.0 | 85.4 | -| StarWAM-CD | 79.0 | 79.2 | 79.1 | -| Streaming-WAM (Ours) | 87.2 | **88.8** | 87.6 | +| π₀ | 65.92 | 58.40 | 62.20 | +| π₀.₅ | 82.74 | 76.76 | 79.80 | +| Motus | 88.66 | 87.02 | 87.80 | +| Motus from WAN2.2 | 77.56 | 77.00 | 77.30 | +| FastWAM-Joint | 86.40 | 87.60 | 87.00 | +| FastWAM-Joint-CD | 86.20 | 85.80 | 86.00 | +| Streaming-WAM (FastWAM, Ours) | **90.40** | **90.80** | **90.60** | +| StarWAM-Joint | 87.80 | 84.60 | 86.20 | +| StarWAM-CD | 85.40 | 86.20 | 85.80 | +| Streaming-WAM (StarWAM, Ours) | 86.60 | 85.80 | 86.20 | #### RoboCasa @@ -164,14 +166,17 @@ Task success alone does not characterize runtime efficiency. We therefore report | LIBERO | Streaming-WAM | 41.0 ms | 5.36 s Long / 3.15 s Short | | LIBERO | Streaming-WAM w/o Action Conditioning | 35.1 ms | 5.20 s Long / 2.92 s Short | | LIBERO | Streaming-WAM w/o Slot Encoder | 36.3 ms | 5.31 s Long / 3.01 s Short | -| RoboTwin 2.0 | StarWAM-Joint | 190.17 ms | 110.22 s | -| RoboTwin 2.0 | StarWAM-CD | 81.21 ms | 102.59 s | -| RoboTwin 2.0 | Streaming-WAM | 47.09 ms | 77.48 s | +| RoboTwin 2.0 | FastWAM-Joint | 652.1 ms | 24.26 s | +| RoboTwin 2.0 | FastWAM-Joint-CD | 165.2 ms | 18.63 s | +| RoboTwin 2.0 | Streaming-WAM (FastWAM) | 54.4 ms | 20.14 s | +| RoboTwin 2.0 | StarWAM-Joint | 196.5 ms | 25.76 s | +| RoboTwin 2.0 | StarWAM-CD | 83.1 ms | 26.23 s | +| RoboTwin 2.0 | Streaming-WAM (StarWAM) | 36.6 ms | 24.44 s | | RoboCasa | X-WAM | 374.07 ms | 17.36 s | | RoboCasa | X-WAM-CD | 134.37 ms | 13.04 s | | RoboCasa | Streaming-WAM | 115.98 ms | 9.49 s | -Across all three benchmarks, Streaming-WAM reduces both runtime measures while maintaining comparable task success. On LIBERO, it achieves a 12.0× Chunk Time speedup over FastWAM and Episode Time speedups of 3.0× and 2.6× on Long and Short tasks, respectively, with 98.20% average success. On RoboTwin 2.0, relative to StarWAM-Joint, Chunk Time falls from 190.17 ms to 47.09 ms and Episode Time from 110.22 s to 77.48 s, while overall success increases from 85.4 to 87.6. On RoboCasa, relative to X-WAM, Streaming-WAM achieves a 3.2× Chunk Time speedup and a 1.8× Episode Time speedup, with comparable average success (75.35% versus 75.42%). +Across all three benchmarks, Streaming-WAM reduces Chunk Time while maintaining competitive task success. On LIBERO, it achieves a 12.0× Chunk Time speedup over FastWAM and Episode Time speedups of 3.0× and 2.6× on Long and Short tasks, respectively, with 98.20% average success. On RoboTwin 2.0, the FastWAM version reduces Chunk Time from 652.1 ms to 54.4 ms and Episode Time from 24.26 s to 20.14 s, while success rises from 87.00 to 90.60. The StarWAM version reduces Chunk Time from 196.5 ms to 36.6 ms and Episode Time from 25.76 s to 24.44 s, with the same 86.20 overall success. On RoboCasa, relative to X-WAM, Streaming-WAM achieves a 3.2× Chunk Time speedup and a 1.8× Episode Time speedup, with comparable average success (75.35% versus 75.42%). ## Runtime layout diff --git a/docs/assets/streaming-wam-chunk-time.png b/docs/assets/streaming-wam-chunk-time.png index d249450..f283c20 100644 Binary files a/docs/assets/streaming-wam-chunk-time.png and b/docs/assets/streaming-wam-chunk-time.png differ diff --git a/docs/assets/streaming-wam-episode-time.png b/docs/assets/streaming-wam-episode-time.png index 8766ce7..d3c611c 100644 Binary files a/docs/assets/streaming-wam-episode-time.png and b/docs/assets/streaming-wam-episode-time.png differ diff --git a/docs/generate_latency_figure.py b/docs/generate_latency_figure.py index 031bd3e..b4c4c66 100644 --- a/docs/generate_latency_figure.py +++ b/docs/generate_latency_figure.py @@ -41,8 +41,8 @@ LIBERO_SHORT = (8.25, 3.74, 3.20, 3.15, 2.92, 3.01) ROBOTWIN_METHODS = ("StarWAM\nJoint", "StarWAM\nCD", "Streaming-\nWAM") -ROBOTWIN_CHUNK = (190.17, 81.21, 47.09) -ROBOTWIN_EPISODE = (110.22, 102.59, 77.48) +ROBOTWIN_CHUNK = (196.5, 83.1, 36.6) +ROBOTWIN_EPISODE = (25.76, 26.23, 24.44) ROBOCASA_METHODS = ("X-WAM", "X-WAM\nCD", "Streaming-\nWAM") ROBOCASA_CHUNK = (374.07, 134.37, 115.98) ROBOCASA_EPISODE = (17.36, 13.04, 9.49) @@ -240,7 +240,7 @@ def render_episode_time(output_path: Path) -> None: _draw_libero_episode(axes[0]) _draw_three_method_panel( axes[1], title="RoboTwin 2.0", ylabel="Seconds", methods=ROBOTWIN_METHODS, - values=ROBOTWIN_EPISODE, ceiling=125, + values=ROBOTWIN_EPISODE, ceiling=30, ) _draw_three_method_panel( axes[2], title="RoboCasa", ylabel="Seconds", methods=ROBOCASA_METHODS, diff --git a/docs/index.html b/docs/index.html index 876208f..acdfa0f 100644 --- a/docs/index.html +++ b/docs/index.html @@ -302,14 +302,16 @@

RoboTwin 2.0

RoboTwin 2.0 clean and randomized results MethodClean ↑Random ↑Total ↑ - π₀65.9258.4062.2 - π₀.₅82.7476.7679.8 - Motus88.6687.0287.8 - Motus from WAN2.277.5677.0077.3 - FastWAM-Joint87.887.3287.56 - StarWAM-Joint84.886.085.4 - StarWAM-CD79.079.279.1 - Streaming-WAM (Ours)87.288.887.6 + π₀65.9258.4062.20 + π₀.₅82.7476.7679.80 + Motus88.6687.0287.80 + Motus from WAN2.277.5677.0077.30 + FastWAM-Joint86.4087.6087.00 + FastWAM-Joint-CD86.2085.8086.00 + Streaming-WAM (FastWAM, Ours)90.4090.8090.60 + StarWAM-Joint87.8084.6086.20 + StarWAM-CD85.4086.2085.80 + Streaming-WAM (StarWAM, Ours)86.6085.8086.20 @@ -357,7 +359,7 @@

Real robot evaluation

Inference efficiency. Task success alone does not characterize runtime efficiency. We therefore report Chunk Time, the latency required to prepare the next action chunk, and Episode Time, the duration of a complete rollout, including inference, execution, and replanning.

-

Across all three benchmarks, Streaming-WAM reduces both runtime measures while maintaining comparable task success. On LIBERO, it achieves a 12.0× Chunk Time speedup over FastWAM and Episode Time speedups of 3.0× and 2.6× on Long and Short tasks, respectively, with 98.20% average success. On RoboTwin 2.0, relative to StarWAM-Joint, Chunk Time falls from 190.17 ms to 47.09 ms and Episode Time from 110.22 s to 77.48 s, while overall success increases from 85.4 to 87.6. On RoboCasa, relative to X-WAM, Streaming-WAM achieves a 3.2× Chunk Time speedup and a 1.8× Episode Time speedup, with comparable average success (75.35% versus 75.42%).

+

Across all three benchmarks, Streaming-WAM reduces Chunk Time while maintaining competitive task success. On LIBERO, it achieves a 12.0× Chunk Time speedup over FastWAM and Episode Time speedups of 3.0× and 2.6× on Long and Short tasks, respectively, with 98.20% average success. On RoboTwin 2.0, the FastWAM version reduces Chunk Time from 652.1 ms to 54.4 ms and Episode Time from 24.26 s to 20.14 s, while success rises from 87.00 to 90.60. The StarWAM version reduces Chunk Time from 196.5 ms to 36.6 ms and Episode Time from 25.76 s to 24.44 s, with the same 86.20 overall success. On RoboCasa, relative to X-WAM, Streaming-WAM achieves a 3.2× Chunk Time speedup and a 1.8× Episode Time speedup, with comparable average success (75.35% versus 75.42%).

@@ -385,9 +387,12 @@

Real robot evaluation

LIBEROStreaming-WAM41.0 ms5.36 s Long / 3.15 s Short LIBEROStreaming-WAM w/o Action Conditioning35.1 ms5.20 s Long / 2.92 s Short LIBEROStreaming-WAM w/o Slot Encoder36.3 ms5.31 s Long / 3.01 s Short - RoboTwin 2.0StarWAM-Joint190.17 ms110.22 s - RoboTwin 2.0StarWAM-CD81.21 ms102.59 s - RoboTwin 2.0Streaming-WAM47.09 ms77.48 s + RoboTwin 2.0FastWAM-Joint652.1 ms24.26 s + RoboTwin 2.0FastWAM-Joint-CD165.2 ms18.63 s + RoboTwin 2.0Streaming-WAM (FastWAM)54.4 ms20.14 s + RoboTwin 2.0StarWAM-Joint196.5 ms25.76 s + RoboTwin 2.0StarWAM-CD83.1 ms26.23 s + RoboTwin 2.0Streaming-WAM (StarWAM)36.6 ms24.44 s RoboCasaX-WAM374.07 ms17.36 s RoboCasaX-WAM-CD134.37 ms13.04 s RoboCasaStreaming-WAM115.98 ms9.49 s diff --git a/examples/robotwin/RoboTwin.md b/examples/robotwin/RoboTwin.md index f8e57e1..a79fc2b 100644 --- a/examples/robotwin/RoboTwin.md +++ b/examples/robotwin/RoboTwin.md @@ -261,7 +261,42 @@ python script/eval_policy.py \ For Shared-DiT, also override `config_path` to the Shared-DiT recipe, use checkpoint-45000, and set both inference-step values to 16. -### 8.3 Results +### 8.3 Canonical cross-family benchmark + +Use the same public harness for StarWAM, FastWAM, and Streaming-WAM checkpoints. +The wrapper defaults to 100 trials per task-setting and `replan_steps=16`; every +path and runtime remains overridable through environment variables. + +```bash +MODE=baseline \ +CHECKPOINT_FORMAT=fastwam \ +CKPT=/path/to/fastwam_joint.pt \ +CONFIG=examples/robotwin/configs/recipes/streamingwam_robotwin_mot_wan22_5b.yaml \ +STATS_PATH=/path/to/action_stats.json \ +BACKBONE_PATH=/path/to/Wan2.2-TI2V-5B \ +ROBOTWIN_HOME=/path/to/RoboTwin \ +INFERENCE_PYTHON=/path/to/streamingwam-env/bin/python \ +SIMULATOR_PYTHON=/path/to/robotwin-env/bin/python \ +bash examples/robotwin/scripts/run_streamingwam_robotwin_benchmark.sh +``` + +The released defaults are selected from `CHECKPOINT_FORMAT` and `MODE`: + +| Checkpoint family | Mode | Video steps | Action steps | +|---|---|---:|---:| +| StarWAM | baseline | 4 | 4 | +| StarWAM | cd | 1 | 1 | +| StarWAM | ac-stream | 1 | 1 | +| FastWAM | baseline | 10 | 10 | +| FastWAM | cd | 1 | 2 | +| Streaming-WAM/FastWAM Stage2 | ac-stream | 1 | 2 | + +Set `NUM_INFERENCE_STEPS` and `ACTION_NUM_INFERENCE_STEPS` to override these +defaults for another checkpoint. Use `AC_STREAM_BACKEND=eager` for the eager +ablation; accelerated AC-Stream is the default. `MODEL_SEED` controls model +noise and `EPISODE_SEED` independently controls the RoboTwin scene chain. + +### 8.4 Results RoboTwin writes one result per (task, config) under `RoboTwin/eval_result/////