Full benchmark results

The complete per-benchmark tables behind OpenWAM-α — the nine simulation leaderboards of Appendix E and the three real-robot suites — with every baseline the source tables report, at the metric each one uses. The home page shows the summary column; this is everything under it.

LIBERO

Evaluation Results on LIBERO. Bold denotes best values, underline second best.

MethodSpatialObjectGoalLongAvg
VLA
OpenVLA84.788.479.253.776.5
π₀98.096.894.488.494.4
StarVLA97.898.696.293.896.6
π₀.₅98.898.298.092.496.9
GR00T-N1.697.798.597.594.497.0
OpenVLA-OFT97.698.497.994.597.1
X-VLA98.298.697.897.698.1
ABot-M098.899.899.096.698.6
Being-H0.599.299.699.497.498.9
Qwen-RobotManip99.2
WAM
Fast-WAM98.2100.097.095.297.6
Motus96.899.896.697.697.7
ImageWAM97.299.298.898.498.4
LingBot-VA98.599.697.298.598.5
DiT4DiT98.6
Being-H0.799.2
ABot-M0.5100.099.899.498.499.4
OpenWAM-α99.699.699.898.299.3

Bold is best, underline second best, as marked in the paper. “—” means the source table does not report that entry.

Transcribed from tables/*.tex in the OpenWAM paper repository.