Welcome to ViLMa, the 2nd Workshop on Visual Localization and Mapping; From Optimization to 3D Foundation Models, organized at ECCV 2026 in Malmö, Sweden.
Topic
Visual localization and 3D mapping stand at a critical inflection point. For decades, the field relied on explicit geometric modeling and hand-crafted optimization. Today, the rapid ascent of data-driven paradigms — specifically 3D foundation models, neural radiance fields (NeRFs), and 3D Gaussian Splatting — challenges these traditional pillars. We are moving from systems that map geometry to models that understand spatial context.
The 2nd Workshop on Visual Localization and Mapping (ViLMa) addresses the central conflict defining this era: the integration of rigorous geometric priors with the generalization capabilities of large-scale learning.
We will tackle the most pressing questions:
- Obsolescence vs. Evolution: Are classical SLAM pipelines becoming obsolete, or must they evolve into the “operating system” for neural representations?
- Hybrid Optimization Pipelines: How do we mathematically fuse probabilistic state estimation with learned implicit features?
- Implicit vs. Explicit: What is the optimal trade-off between memory-efficient implicit representations and actionable explicit maps?
ViLMa covers the frontier of Spatial AI: next-generation SLAM, hybrid neuro-geometric pipelines, open-vocabulary mapping, and world models for embodied agents. By convening leading experts from academia and industry, we aim to foster a deeper synthesis between geometric principles and modern deep learning paradigms. This workshop is not just a review of current methods; it is a forum to define the architecture of future perception systems for AR/VR, robotics, and autonomous driving.
Topics of Interest
- Next-generation SLAM: real-time localization and mapping built on learned representations, and what endures from the classical pipeline.
- 3D foundation models: feed-forward reconstruction, pose estimation and correspondence from large-scale pre-training, and the limits of generalization.
- Neural and Gaussian scene representations: NeRFs, 3D Gaussian Splatting and their successors used as maps rather than as renderers.
- Hybrid neuro-geometric optimization: fusing probabilistic state estimation and bundle adjustment with learned features and priors.
- Implicit versus explicit maps: memory footprint, queryability, and the trade-off between compact implicit and directly actionable explicit representations.
- Open-vocabulary and semantic mapping: language-grounded maps and open-set scene understanding for downstream tasks.
- World models for embodied agents: predictive spatial models that support navigation, manipulation and planning.
- Large-scale and lifelong mapping: city-scale reconstruction, map maintenance, change detection and long-term operation.
- Sensor fusion: combining cameras with LiDAR, IMU, radar and GNSS inside learned pipelines.
- Robustness and generalization: low light and adverse weather, dynamic scenes, occlusion, and domain shift.
- Benchmarks and evaluation: datasets, metrics and protocols that can assess learned and hybrid systems fairly.
- Applications and deployment: AR/VR, robotics and autonomous driving, including latency, efficiency and on-device constraints.
Schedule
- Tuesday, September 8, 2026, 08:00 - 12:00 — ViLMa Workshop @ ECCV 2026
- Room: TBA
NOTE: Times are shown in Central European Summer Time (UTC+2). The programme is preliminary and subject to change.
| 08:00 - 08:10 | Introduction to the Workshop |
| 08:10 - 08:50 | Invited Talk 1 |
| 08:50 - 09:30 | Invited Talk 2 |
| 09:30 - 10:00 | Coffee Break (30 min) |
| 10:00 - 10:40 | Invited Talk 3 |
| 10:40 - 11:20 | Invited Talk 4 |
| 11:20 - 11:50 | Panel Discussion |
| 11:50 - 12:00 | Closing Remarks |
Invited Speakers
Martin R. Oswald
Assistant Professor
University of Amsterdam
Paul-Edouard Sarlin
Researcher
Andrea Vedaldi
Professor
University of Oxford
Christian Wolf
Principal Scientist
Naver Labs Europe
Martin R. Oswald is an Assistant Professor in the Computer Vision Group at the University of Amsterdam and a senior researcher in the Computer Vision and Geometry lab at ETH Zurich, where he was previously a postdoctoral researcher with Marc Pollefeys. He obtained his PhD in computer vision from the Technical University of Munich in 2015, advised by Daniel Cremers, after a Diplom in computer science from TU Dresden and a master’s degree in civil engineering from Universidad Técnica Federico Santa María in Valparaíso, Chile. His research covers 3D reconstruction, 3D scene understanding and SLAM. He has co-authored a series of widely used neural SLAM systems, among them NICE-SLAM (CVPR 2022), Point-SLAM (ICCV 2023) and NICER-SLAM, which received a Best Paper Honorable Mention at 3DV 2024.
Paul-Edouard Sarlin is a researcher at Google in Zurich, where he works on world-scale mapping. He completed his PhD in 2024 with the Computer Vision and Geometry group at ETH Zurich under the supervision of Marc Pollefeys; his thesis, On Learning and Geometry for Visual Localization and Mapping, was awarded the ETH Silver Medal and the DAGM MVTec Dissertation Award at GCPR 2025. He is best known for the learned feature matchers SuperGlue (CVPR 2020) and LightGlue (ICCV 2023), and for hloc, a widely used open-source toolbox for visual localization and structure-from-motion. His further work includes Pixel-Perfect Structure-from-Motion (ICCV 2021), the LaMAR benchmark for AR localization and mapping, and OrienterNet, SNAP and GeoCalib.
Andrea Vedaldi is Professor of Computer Vision and Machine Learning at the University of Oxford, where he co-leads the Visual Geometry Group (VGG) in the Department of Engineering Science. His research develops methods that understand the content of images and videos automatically, with little to no manual supervision, both in terms of semantics and of 3D geometry. He was elected a Fellow of the Royal Academy of Engineering in 2025 and is one of the inaugural recipients of the Royal Society Faraday Discovery Fellowship, which supports a long-term programme on spatial artificial intelligence. His recent work includes the VGGT family of feed-forward 3D reconstruction models, whose quality scales with data and model size; VGGT-Ω was named a best paper finalist at CVPR 2026.
Christian Wolf is a Principal Scientist at Naver Labs Europe, where he leads the Spatial AI team. From 2005 to 2021 he was an associate professor (Maître de Conférences, HDR) at INSA Lyon and the CNRS laboratory LIRIS, where he held the ANR/Naver/INSA chair in artificial intelligence “REMEMBER — Learning Reasoning, Memory and Behavior”. He received his MSc from TU Vienna in 2000, his PhD from INSA Lyon in 2003, and his habilitation in 2012. His research focuses on AI for robotics, in particular machine learning and embodied computer vision, the large-scale learning of high-level reasoning from visual observations, and the connections between machine learning and control. He is an ELLIS member and served as an associate editor of IEEE TPAMI from 2019 to 2025.
Organizers
Luca Carlone
Associate Professor
Massachusetts Institute of Technology
Daniel Cremers
Professor
Technical University of Munich
Dima Damen
Professor
University of Bristol
Frank Dellaert
Professor
Georgia Institute of Technology
Ayoung Kim
Professor
Seoul National University
Patrick Wenzel
AI Research Engineer
Helsing
Niclas Zeller
Professor
Karlsruhe University of Applied Sciences
Related Workshops
Other ECCV 2026 workshops on closely related topics:
- 3rd Neural SLAM Workshop (NeuSLAM) — September 8, afternoon
- Privacy-Preserving Visual Localization and Mapping — September 8, afternoon
- Structure-from-Motion in the Age of Deep Learning (SfM-ADL) — September 9, morning
ViLMa takes place on the morning of September 8, so none of these clash with our programme.