A Linear Probe Reads Task Progress Out of π0.5
Vision-language-action models are being treated like deployable workers. We still have almost no cheap way to ask one “how far along are you?” Cornell researchers probe π0.5 and find that task progress — normalized time remaining in a successful trajectory — is linearly readable from the residual stream.
The paper landed on arXiv August 13. Lead authors Atiksh Bhardwaj and Edward Weiyi Duan share credit; Wei-Chiu Ma and Preston Culbertson are on the line.
Readable, not steerable
They separate three claims people blur together:
- Weak decodability: a linear probe tracks progress on the training distribution
- Strong decodability: the same probe still tracks when you swap the language prompt
- Steerability: injecting the feature changes what the policy does
Progress clears the first two. It does not clear the third. You can read the clock. You cannot twist it.
The signal is already in the pretrained PaliGemma backbone, before any robot data (mean R² = 0.700, MAE 0.123 on the paper’s progress plot). After robot training it tightens: base π0.5 hits R² = 0.926, fine-tuned π0.5 0.927. A single probe trained on multi-prompt data generalizes to unseen tasks and moves under language counterfactuals. Fine-tuning still dulls language sensitivity, and the probe is used as a diagnostic for that fade.
A stall detector you did not have to label
The practical use is a label-free out-of-distribution detector. If predicted progress stops moving, treat the rollout as stuck. The authors say that detector is competitive with supervised VLA monitors such as SAFE, including on held-out tasks and held-out perturbation modes.
They are honest about the definition: progress is normalized time on successful demonstrations, so it mixes “the clock ran out” with “the mug is in the sink.” That is a limit, not a footnote.
A Human’s Take
I want a cheap heartbeat on every deployed VLA, and leftover-time is a better heartbeat than “the softmax looks weird.” The no-steering result is the adult part of the paper. If you can read a feature and cannot write it, do not sell an activation-edit as a recovery button. Wire it as a monitor, abort the stall, and collect the failure.