Papers
arxiv:2608.16885

τ_0-VLA: a Hierarchical Robot Foundation Model with World-Model-Guided Test-Time Computation

Published on Aug 17
· Submitted by
jrryzh(SII)
on Aug 21
Authors:
,
,
,
,
,
,
,
,
,
,
,
,
,
,
,
,
,
,
,
,
,

Abstract

A hierarchical vision-language-action model improves long-horizon robot manipulation by using world-model-guided test-time search to scale computation for high-level subtask decisions.

Long-horizon robot manipulation requires a robot to both execute individual skills reliably and sequence them coherently over extended tasks. Most hierarchical vision-language-action (VLA) models make each such decision with a single forward pass, leaving no mechanism to allocate additional computation to difficult or consequential choices. We introduce τ_0-VLA, a hierarchical robot foundation model that formulates high-level subtask generation as a compute-scalable inference problem through world-model-guided test-time computation. At each inference step, the high-level policy uses execution memory to generate a subtask and, when needed, searches over alternatives before committing to its output. A low-level policy then executes the generated subtask across multiple robot embodiments. The policy is trained on 40,115 hours of heterogeneous real-world data with multimodal co-training. Across in-domain and distribution-shifted settings, allocating additional test-time computation substantially improves next-subtask prediction accuracy, and these gains translate into higher closed-loop success on long-horizon robot manipulation tasks.

Community

Paper author Paper submitter

🤖 What if a robot could compare possible futures before deciding what to do next?

We introduce τ₀-VLA, a hierarchical robot foundation model for long-horizon manipulation. Its high-level policy maintains execution memory and, when a decision is uncertain, allocates additional test-time computation to propose candidate subtasks, predict their visual consequences with a world model, and compare alternatives before committing. A generalist low-level VLA then executes the selected subtask across robot embodiments.

Highlights:

  • The low-level policy is trained on 40,115 hours of heterogeneous real-world robot data with multimodal co-training.
  • Selective test-time computation improves next-subtask prediction accuracy by 15–24 percentage points across in-domain and distribution-shifted settings.
  • We evaluate real-world manipulation tasks containing 13–25 ordered steps, with episodes lasting up to 12 minutes.
  • Using the same low-level policy, hierarchical planning improves average closed-loop success from 27.5% to 45.0% across four long-horizon tasks.
  • We release the official code and pretrained low-level VLA checkpoint, with high-level policy on the way.

🌐 Project page
💻 Code
🤗 Model checkpoint

Questions and feedback are very welcome!

This is an automated message from the Librarian Bot. I found the following papers similar to this paper.

The following papers were recommended by the Semantic Scholar API

Please give a thumbs up to this comment if you found it helpful!

If you want recommendations for any Paper on Hugging Face checkout this Space

You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: @librarian-bot recommend

Sign up or log in to comment

Get this paper in your agent:

hf papers read 2608.16885
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper 1

Datasets citing this paper 0

No dataset linking this paper

Cite arxiv.org/abs/2608.16885 in a dataset README.md to link it from this page.

Spaces citing this paper 0

No Space linking this paper

Cite arxiv.org/abs/2608.16885 in a Space README.md to link it from this page.

Collections including this paper 0

No Collection including this paper

Add this paper to a collection to link it from this page.
MiniMax H3 Video Generator 20 free credits · Text & image to video Try Free →