With the initial state held fixed, the flow sampler, not the environment, carries
61 to 84 percent of the variance in a reported success rate. Nobody had decomposed
robot policy evaluation variance by source; once you do, the budget follows. A
replicate block splits that surviving share on the informative task into the flow
sampler (38.5 percent of total) and system nondeterminism (46.0 percent), the second
of which is invisible in any single reported rate. Rerunning the
field's standard 500-episode LIBERO protocol changing only the sampling seed moves the
headline number by 2.2 points, a band already containing 27 percent of the recomputable
improvements claimed across 50 recent VLA papers, 34.4 percent of whose highlighted
comparisons state no usable rollout count at all.
Where the variance lives, and what it costs. Top: with the initial state held
fixed, the flow sampler accounts for most outcome variance across three policies. Bottom:
what each factor moves on the reported rate. The seed shifts the 500-episode aggregate by
only 2.2 points precisely because averaging suppresses the per-episode variance above it,
and that gap is the budget an honest protocol has to buy.The whole Pareto frontier is one cell: a single denoising step executed
over a horizon of 10. The published default (10, 5) is dominated.
Findings
The scene is the smallest part of the variance. With the initial state held
fixed, 84, 81, and 61 percent of outcome variance survives for pi0.5, pi0, and
SmolVLA, across base rates from 0.16 to 0.59, replicating at three environment
seeds. A replicate block splits that surviving share on the informative task into
the flow sampler (38.5 percent of total) and system nondeterminism (46.0 percent).
That the sampler can decide an episode is established; the partition itself had not
been measured for robot policies, and it is what tells you how many rollouts a claim
actually needs.
The budget that follows. Rerunning the field's standard 500-episode LIBERO
protocol changing only the sampling seed moves the headline number by 2.2 points, a
band already containing 27 percent of the recomputable improvements claimed across 50
recent VLA papers. At the audited median of 50 rollouts per task the minimum
detectable effect is 18 points; a 2-point claim costs 2,200 episodes per arm.
Spending it: the shipped default is dominated. For pi0.5 a single Euler step at
horizon 10 matches the 10-step default at 2.7x lower measured inference latency,
statistically indistinguishable under a pre-declared equivalence margin, on the
unmodified public checkpoint. Its one confirmed price is 4.7 points under
camera-viewpoint shift (pre-declared 838-pair sample, exact McNemar p=0.0008). The
equivalence certificate is scoped honestly: it holds on the registered battery, which
is near ceiling, and the study says where it stops holding. Others have shown few
steps can suffice and that the two knobs couple; what the powered factorial on
released weights adds is a claim of equivalence rather than an anecdote.
Rankings survive off the floor; margins never do. Across four LIBERO-Plus axes,
one 16.5-point clean gap stands for anything from 9.6 to 67 points of deployed margin
(the clean anchors are single uncontrolled-seed draws, disclosed with their seed
band in the paper).
The single significant rank inversion (SmolVLA over pi0 under robot-state shift, exact
p=0.0022; exploratory, post-freeze axis) appears exactly where an axis floors the
trailing policies.
A third of published claims cannot be checked at all. Across 50 audited VLA
papers, 34.4 percent of highlighted comparisons state no usable rollout count, so the
claim cannot be evaluated even in principle.
Two perturbation axes, three policies, frozen variant sets. Lighting degrades
without reordering; robot-initial-state shift floors pi0, where the clean pi0 versus SmolVLA
ordering inverts (exact p=0.0022; exploratory, post-freeze axis).
Tooling
Every number regenerates from released per-episode JSONL records by one script, including the
tables and the abstract's numerals. The harness is pip-installable
(pip install roborigor): exact intervals, paired comparisons, a rollout-budget
calculator, variance decomposition, and a campaign runner with exactly-once resume.
BibTeX
Citation will be posted when the preprint is up (in progress).
Page assets and numbers are generated from the released data; see the repository
for the drift test that keeps them honest.