This is an outdated version published on 2026-08-16. Read the most recent version.
Preprint / Version 1

Auditing Manipulation Benchmarks Before Training

Placement Radius, a Shipped-State Defect, and Embodiment Tolerance

##article.authors##

  • Avneh Bhatia Exea Labs https://orcid.org/0009-0006-7072-4014
  • Monish Allada University of North Carolina at Charlotte
  • Lalith Samineni University of Maryland
  • Joshua Selvaraj Exea Labs

DOI:

https://doi.org/10.31224/7955

Keywords:

VLA, AI Auditability, Auditing, Machine Learning, Explainable AI, Counterfactual Explanations, Anchor Explanations, Feature Sensitivity, Active Learning, Model Interpretability, Fairness, Membership Query Learning

Abstract

Manipulation benchmarks are often interpreted as tests of perception and language-conditioned control, but a score can also reflect regularities in the evaluation states. We introduce the placement radius R, the radius of the smallest disc containing an object’s shipped positions. The condition R ≤δ is exactly the geometric condition for one task-indexed constant to lie within distance δ of every shipped position; behavioral equivalence additionally requires an embodiment-specific robustness assumption. Applied to LIBERO, this audit reveals a shipped-state defect in LIBERO-Object. Its tasks declare 5×5 cm target-sampling regions, yet the shipped states span only 0.0–1.9 cm, and six of ten tasks place the target at one point to floating-point precision. In a paired intervention on one task, sampling the declared region reduces the lookup-goal condition from 25/30 to 10/30 successes (p = 0.0007). A reset-informed condition supplied the true target position and succeeded on all 30 repaired states. Thus the same motion controller can solve the repaired states when supplied the correct target. A sparse displacement ladder shows a dose response but does not identify a uniform tolerance δ: outcomes change even at the smallest nonzero level, and the experiment varies direction and initial state together. Accordingly, placement radius is a pre-training diagnostic of geometric shortcut opportunity, not by itself a certificate of policy behavior. The causal intervention is limited to one task, and we do not establish that a trained policy exploits the defect.

Downloads

Download data is not yet available.

Downloads

Posted

2026-08-16

Versions