Video language planning

Y Du, S Yang, P Florence, F Xia, A Wahid… - International …, 2024 - proceedings.iclr.cc
… to integrate vision-language models and text-to-video models to enable video language
planning (VLP), where given the current image observation and a language instruction, the …

Navid: Video-based vlm plans the next step for vision-and-language navigation

J Zhang, K Wang, R Xu, G Zhou, Y Hong… - arXiv preprint arXiv …, 2024 - arxiv.org
… NaVid, a video-based large vision language model (VLM), to … an on-the-fly video stream
from a monocular RGB camera … Moreover, our video-based approach can effectively encode …

This&that: Language-gesture controlled video generation for robot planning

B Wang, N Sridhar, C Feng… - … on Robotics and …, 2025 - ieeexplore.ieee.org
… We consider the video generator as a generalizable planner that envisions how … video
generation method that combines language and gesture instructions to create robot action plans

Plan-X: Instruct video generation via semantic planning

L Huang, Y Xie, H Xu, T Gu, C Zhang, G Song… - arXiv preprint arXiv …, 2025 - arxiv.org
… We extensively evaluate Plan-X on a challenging video generation benchmark covering
text-to-video, image-tovideo, and video continuation tasks. Across all settings, Plan-X achieves …

NovaPlan: Zero-shot long-horizon manipulation via closed-loop video language planning

J Fu, J Nan, L Sun, H Li, J Qian, JL Barry… - arXiv preprint arXiv …, 2026 - arxiv.org
… While vision-language models (VLMs) and videovideo planning with geometrically grounded
robot execution for zero-shot long-horizon manipulation. At the high level, a VLM planner

Vg-tvp: Multimodal procedural planning via visually grounded text-video prompting

MF Ilaslan, A Köksal, KQ Lin, B Satar… - Proceedings of the …, 2025 - ojs.aaai.org
… between text and video plans) and spatial consistency (ie subsequent video steps must …
We follow this alignment to generate the relevant task videos which compose the Video Plan

Lamp: Language-assisted motion planning for controllable video generation

MB Kizil, E Sanli, NJ Mitra, E Erdem… - Proceedings of the …, 2026 - openaccess.thecvf.com
… and camera motion planning in a shared 3D space, a key requirement for cinematographically
coherent video synthesis (see Fig. 1). To train the LLM motion planner, we construct a …

See, plan, predict: Language-guided cognitive planning with video prediction

M Attarian, A Gupta, Z Zhou, W Yu… - arXiv preprint arXiv …, 2022 - arxiv.org
… Our language planner generated the following series of instructions: ”Pick and place the
purple F on the rightmost quarter,” ”Pick and place the green R on the left of the purple F ,” ”Pick …

VideoAgent: Long-Form Video Understanding with Large Language Model as Agent

X Wang, Y Zhang, O Zohar, S Yeung-Levy - European Conference on …, 2024 - Springer
… Long-form video understanding represents a significant challenge within computer vision, …
for long-form video understanding, we emphasize interactive reasoning and planning over the …

Videodirectorgpt: Consistent multi-scene video generation via llm-guided planning

H Lin, A Zala, J Cho, M Bansal - arXiv preprint arXiv:2309.15091, 2023 - arxiv.org
planning and grounded video generation. Specifically, given a single text prompt, we first ask
our video planner LLM (GPT-4) to expand it into a ‘video plan’… this video plan, our video gen…