Open-o3-Video: Grounded Video Reasoning with Explicit Spatio-Temporal Evidence

Jul 1, 2026·
J Meng
,
X Li
,
H Wang
Yue Tan
Yue Tan
,
T Zhang
,
L Kong
,
Y Tong
,
A Wang
,
Z Teng
,
et al.
· 1 min read
Type
Publication
Proceedings of the 43rd International Conference on Machine Learning (ICML 2026)
Status
Peer-reviewed Open access
publications

This publication entry was migrated from the previous site.

Citation: J Meng, X Li, H Wang, Yue Tan, T Zhang, L Kong, Y Tong, A Wang, Z Teng, et al. (2026). “Open-o3-Video: Grounded Video Reasoning with Explicit Spatio-Temporal Evidence.” ICML 2026.

Yue Tan
Authors
Yue Tan received a BS in Information and Computing Science from the School of Information Science and Technology, Peking University (2022–2026). They are a Ph.D. student since 2026 at the Wang Xuan Institute of Computer Technology, School of Computer Science, Peking University. Research interests include multimodal large models, video understanding, and AIGC.