Full benchmark results
The complete per-benchmark tables behind OpenWAM-α — the nine simulation leaderboards of Appendix E and the three real-robot suites — with every baseline the source tables report, at the metric each one uses. The home page shows the summary column; this is everything under it.
LIBERO
Evaluation Results on LIBERO. Bold denotes best values, underline second best.
| Method | Spatial | Object | Goal | Long | Avg |
|---|---|---|---|---|---|
| VLA | |||||
| OpenVLA | 84.7 | 88.4 | 79.2 | 53.7 | 76.5 |
| π₀ | 98.0 | 96.8 | 94.4 | 88.4 | 94.4 |
| StarVLA | 97.8 | 98.6 | 96.2 | 93.8 | 96.6 |
| π₀.₅ | 98.8 | 98.2 | 98.0 | 92.4 | 96.9 |
| GR00T-N1.6 | 97.7 | 98.5 | 97.5 | 94.4 | 97.0 |
| OpenVLA-OFT | 97.6 | 98.4 | 97.9 | 94.5 | 97.1 |
| X-VLA | 98.2 | 98.6 | 97.8 | 97.6 | 98.1 |
| ABot-M0 | 98.8 | 99.8 | 99.0 | 96.6 | 98.6 |
| Being-H0.5 | 99.2 | 99.6 | 99.4 | 97.4 | 98.9 |
| Qwen-RobotManip | — | — | — | — | 99.2 |
| WAM | |||||
| Fast-WAM | 98.2 | 100.0 | 97.0 | 95.2 | 97.6 |
| Motus | 96.8 | 99.8 | 96.6 | 97.6 | 97.7 |
| ImageWAM | 97.2 | 99.2 | 98.8 | 98.4 | 98.4 |
| LingBot-VA | 98.5 | 99.6 | 97.2 | 98.5 | 98.5 |
| DiT4DiT | — | — | — | — | 98.6 |
| Being-H0.7 | — | — | — | — | 99.2 |
| ABot-M0.5 | 100.0 | 99.8 | 99.4 | 98.4 | 99.4 |
| OpenWAM-α | 99.6 | 99.6 | 99.8 | 98.2 | 99.3 |
Bold is best, underline second best, as marked in the paper. “—” means the source table does not report that entry.
Transcribed from tables/*.tex in the OpenWAM paper repository.