Program
Date: 9th September (afternoon), ECCV2026, Malmo, Sweden
Room: Quality View Oresundslokalen – 3 (C)
Visual Object Tracking and Segmentation challenge VOTS2026 is the 14th annual benchmarking event organized by the VOT initiative. Four sub-challenges dedicated to different tracking problems are considered. The VOTS2026 subchallenge addresses tracking by video object segmentation, VOTSt2026 addresses tracking transforming objects, VOTSr2026 addresses referring expression tracking, and VOTSp2026 addresses point tracking. The VOTS2026 workshop program reviews the results, the winner methods and features keynotes of prominent researchers in video understanding.
13:45 - 15:30 Session 1 (Chair: Pavel Tokmakov)
- [13:45] Hello from the organizers
- [13:50] Keynote 1: Fatma Guney
- [14:20] VOTS2026 results & insights, Matej Kristan
- [14:40] VOTSp2026 results & insights, Adam Harley
- [15:00] Keynote 2: Amir Bar
15:30 - 16:00 Coffee Break
16:00 - 18:00 Session 2 (Chair: Jiri Matas)
- [16:00] Keynote 3: Mehdi S. M. Sajjadi
- [16:30] VOTSt2026 + VOTSr2026 results & insights, Pavel Tokmakov
- [16:55] VOTSt2026 winners: “MUSMEM – Associating Visual Objects Through Transformers”, Elham Soltani Kazemi, Gani Rahmon, Imad Eddine Toubal, Juan Mogollon, Kannappan Palaniappan, University of Missouri
- [17:10] VOTSr2026 winners: “ReflexTrack – Feedback-Driven Referring Video Segmentation”, Yuanjia Li, Tianyang Xu, Zhangyong Tang, He Wang, Xue-Feng Zhu, Xiao-Jun Wu, Josef Kittler; Jiangnan University & University of Surrey
- [17:25] Keynote 4: Kristen Grauman
- [17:55] Concluding remarks
Keynotes
Prof. Güney is an Assistant Professor at Koç University, where she leads the Autonomous Vision Group. Her research focuses on computer vision for autonomous systems, spanning scene geometry and motion, future prediction, autonomous driving, and point tracking. Among her recognitions are an ERC Starting Grant, a Newton Fund Advanced Fellowship, a Marie Curie Individual Fellowship, and the 2026 Kadir Has Promising Scientist Award.
Prof. Amir Bar is an Assistant Professor at Imperial College London and a Founding Member of Technical Staff at AMI Labs. His research aims to enable machines to perceive, reason, and act from visual data, with recent work centered on self-supervised visual learning, video and embodied world models, and planning. His work has received several distinctions, including a CVPR 2025 Best Paper Honorable Mention for Navigation World Models and an Outstanding Paper Award at the ICLR 2026 World Models Workshop for EB-JEPA.
Dr. Sajjadi is a Research Scientist, Tech Lead & Manager at Google DeepMind. His research focuses on spatio-temporal scene understanding, deep generative models, dynamic 3D/4D scene representations, and visual tracking, with the broader goal of building systems that can perceive and generate the world in 3D. Among the recent highlights of his work is D4RT, which received the CVPR 2026 Best Paper Award.
Prof. Grauman is a Professor and Truchard Family Chair in Natural Sciences at the University of Texas at Austin, where she leads the Computer Vision Group. Her research spans video understanding, egocentric vision, multimodal and embodied perception, with major recent efforts including the large-scale Ego4D and Ego-Exo4D initiatives for understanding human activity from first- and third-person video. Among her many honors are Fellowships of IEEE, AAAS, and AAAI, the IJCAI Computers and Thought Award, the Marr Prize, and the Helmholtz Prize.
