Open-o3-Video: Grounded Video Reasoning with Explicit Spatio-Temporal Evidence
Jul 1, 2026·,,
,,,,,,·
1 min read
J Meng
X Li
H Wang
Yue Tan
T Zhang
L Kong
Y Tong
A Wang
Z Teng
et al.
Type
Publication
Proceedings of the 43rd International Conference on Machine Learning (ICML 2026)
Status
Peer-reviewed
Open access
This publication entry was migrated from the previous site.
Citation: J Meng, X Li, H Wang, Yue Tan, T Zhang, L Kong, Y Tong, A Wang, Z Teng, et al. (2026). “Open-o3-Video: Grounded Video Reasoning with Explicit Spatio-Temporal Evidence.” ICML 2026.

Authors
Yue Tan received a BS in Information and Computing Science from the School of
Information Science and Technology, Peking University (2022–2026). They are a
Ph.D. student since 2026 at the Wang Xuan Institute of Computer Technology,
School of Computer Science, Peking University. Research interests include
multimodal large models, video understanding, and AIGC.