Robotics Institute, Carnegie Mellon University
LLM-based long-horizon task decomposition enables structured demonstration retrieval, which effectively leverages existing skills from prior datasets for data-efficient VLA adaptation.
HSR identifies existing skills in the prior dataset based on task-instruction clustering. Given a
long-horizon target task,
1. HSR first decomposes the task into several subtasks, evaluates and selects the plan with the highest
score;
2. For each subtask, HSR retrieves data based on the subtask instruction, and further reranks using
representative demonstrations from the target task to filter out low-consistency data;
3. HSR first pretrains the policy on retrieved data to learn general skills, and then finetunes it on
task-related data to adapt to the target task.
Close the drawer
Pick up the crumpled paper and throw it in the trashcan
Put the cup in the drawer and close the drawer
Put the tea bag next to the cup and pour water into the cup
Put both the alphabet soup and the tomato sauce in the basket
@misc{hao2026hierarchicalskillretrievaldataefficient,
title={Hierarchical Skill Retrieval for Data-Efficient Adaptation of Vision-Language-Action Models},
author={Haoran Hao and Shahram Najam Syed and Jeff Schneider and Jeffrey Ichnowski},
year={2026},
eprint={2608.24042},
archivePrefix={arXiv},
primaryClass={cs.RO},
url={https://arxiv.org/abs/2608.24042},
}