Auditing Manipulation Benchmarks Before Training
Placement Radius, a Shipped-State Defect, and Embodiment Tolerance
DOI:
https://doi.org/10.31224/7955Keywords:
VLA, AI Auditability, Auditing, Machine Learning, Explainable AI, Counterfactual Explanations, Anchor Explanations, Feature Sensitivity, Active Learning, Model Interpretability, Fairness, Membership Query LearningAbstract
Manipulation benchmarks are often interpreted as tests of perception and language-conditioned control, but a score can also reflect regularities in the evaluation states. We introduce the placement radius R, the radius of the smallest disc containing an object’s shipped positions. The condition R ≤δ is exactly the geometric condition for one task-indexed constant to lie within distance δ of every shipped position; behavioral equivalence additionally requires an embodiment-specific robustness assumption. Applied to LIBERO, this audit reveals a shipped-state defect in LIBERO-Object. Its tasks declare 5×5 cm target-sampling regions, yet the shipped states span only 0.0–1.9 cm, and six of ten tasks place the target at one point to floating-point precision. In a paired intervention on one task, sampling the declared region reduces the lookup-goal condition from 25/30 to 10/30 successes (p = 0.0007). A reset-informed condition supplied the true target position and succeeded on all 30 repaired states. Thus the same motion controller can solve the repaired states when supplied the correct target. A sparse displacement ladder shows a dose response but does not identify a uniform tolerance δ: outcomes change even at the smallest nonzero level, and the experiment varies direction and initial state together. Accordingly, placement radius is a pre-training diagnostic of geometric shortcut opportunity, not by itself a certificate of policy behavior. The causal intervention is limited to one task, and we do not establish that a trained policy exploits the defect.
Downloads
Downloads
Posted
Versions
- 2026-08-17 (2)
- 2026-08-16 (1)
License
Copyright (c) 2026 Avneh Bhatia, Monish Allada, Lalith Samineni, Joshua Selvaraj

This work is licensed under a Creative Commons Attribution 4.0 International License.