Welcome to ViLMa, the 2nd Workshop on Visual Localization and Mapping; From Optimization to 3D Foundation Models, organized at ECCV 2026 in Malmö, Sweden.
Topic
Visual localization and 3D mapping stand at a critical inflection point. For decades, the field relied on explicit geometric modeling and hand-crafted optimization. Today, the rapid ascent of data-driven paradigms — specifically 3D foundation models, neural radiance fields (NeRFs), and 3D Gaussian Splatting — challenges these traditional pillars. We are moving from systems that map geometry to models that understand spatial context.
The 2nd Workshop on Visual Localization and Mapping (ViLMa) addresses the central conflict defining this era: the integration of rigorous geometric priors with the generalization capabilities of large-scale learning.
We will tackle the most pressing questions:
- Obsolescence vs. Evolution: Are classical SLAM pipelines becoming obsolete, or must they evolve into the “operating system” for neural representations?
- Hybrid Optimization Pipelines: How do we mathematically fuse probabilistic state estimation with learned implicit features?
- Implicit vs. Explicit: What is the optimal trade-off between memory-efficient implicit representations and actionable explicit maps?
ViLMa covers the frontier of Spatial AI: next-generation SLAM, hybrid neuro-geometric pipelines, open-vocabulary mapping, and world models for embodied agents. By convening leading experts from academia and industry, we aim to foster a deeper synthesis between geometric principles and modern deep learning paradigms. This workshop is not just a review of current methods; it is a forum to define the architecture of future perception systems for AR/VR, robotics, and autonomous driving.
Topics of Interest
- Next-generation SLAM: real-time localization and mapping built on learned representations, and what endures from the classical pipeline.
- 3D foundation models: feed-forward reconstruction, pose estimation and correspondence from large-scale pre-training, and the limits of generalization.
- Neural and Gaussian scene representations: NeRFs, 3D Gaussian Splatting and their successors used as maps rather than as renderers.
- Hybrid neuro-geometric optimization: fusing probabilistic state estimation and bundle adjustment with learned features and priors.
- Implicit versus explicit maps: memory footprint, queryability, and the trade-off between compact implicit and directly actionable explicit representations.
- Open-vocabulary and semantic mapping: language-grounded maps and open-set scene understanding for downstream tasks.
- World models for embodied agents: predictive spatial models that support navigation, manipulation and planning.
- Large-scale and lifelong mapping: city-scale reconstruction, map maintenance, change detection and long-term operation.
- Sensor fusion: combining cameras with LiDAR, IMU, radar and GNSS inside learned pipelines.
- Robustness and generalization: low light and adverse weather, dynamic scenes, occlusion, and domain shift.
- Benchmarks and evaluation: datasets, metrics and protocols that can assess learned and hybrid systems fairly.
- Applications and deployment: AR/VR, robotics and autonomous driving, including latency, efficiency and on-device constraints.
Schedule
- Tuesday, September 8, 2026, 08:10 - 12:30 — ViLMa Workshop @ ECCV 2026
- Room: Malmömässan K2
NOTE: Times are shown in Central European Summer Time (UTC+2). The programme is preliminary and subject to change.
| 08:10 - 08:15 | Introduction to the Workshop |
| 08:15 - 08:45 | Keynote: Christian Wolf |
| 08:45 - 09:15 | Keynote: Paul-Edouard Sarlin |
| 09:15 - 09:45 | Keynote: Patrick Wenzel |
| 09:45 - 10:15 | Keynote: Andrea Vedaldi |
| 10:15 - 11:00 | Coffee Break (45 min) |
| 11:00 - 11:30 | Keynote: Martin R. Oswald |
| 11:30 - 12:00 | Keynote: Daniel Cremers |
| 12:00 - 12:30 | Panel Discussion & Closing Remarks |
Keynote Speakers
Daniel Cremers
Professor
Technical University of Munich
Martin R. Oswald
Assistant Professor
University of Amsterdam
Paul-Edouard Sarlin
Researcher
Andrea Vedaldi
Professor
University of Oxford
Patrick Wenzel
AI Research Engineer
Helsing
Christian Wolf
Principal Scientist
Naver Labs Europe
Daniel Cremers is Professor of Computer Science and Mathematics at the Technical University of Munich, where he holds the Chair of Computer Vision and Artificial Intelligence, and is a director of the Munich Center for Machine Learning. He obtained his PhD in computer science in 2002, after diplomas in mathematics and physics, and worked as a postdoctoral researcher at UCLA and Siemens Corporate Research before becoming a professor at the University of Bonn in 2005 and moving to TUM in 2009. His research spans variational methods, convex optimization and deep learning for 3D reconstruction, motion estimation and SLAM, including the widely used direct methods LSD-SLAM and DSO. He received the Gottfried Wilhelm Leibniz Award in 2016, Germany’s most prestigious research prize, as well as an ERC Starting, Consolidator and Advanced Grant.
Martin R. Oswald is an Assistant Professor in the Computer Vision Group at the University of Amsterdam and a senior researcher in the Computer Vision and Geometry lab at ETH Zurich, where he was previously a postdoctoral researcher with Marc Pollefeys. He obtained his PhD in computer vision from the Technical University of Munich in 2015, advised by Daniel Cremers, after a Diplom in computer science from TU Dresden and a master’s degree in civil engineering from Universidad Técnica Federico Santa María in Valparaíso, Chile. His research covers 3D reconstruction, 3D scene understanding and SLAM. He has co-authored a series of widely used neural SLAM systems, among them NICE-SLAM (CVPR 2022), Point-SLAM (ICCV 2023) and NICER-SLAM, which received a Best Paper Honorable Mention at 3DV 2024.
Paul-Edouard Sarlin is a researcher at Google in Zurich, where he works on world-scale mapping. He completed his PhD in 2024 with the Computer Vision and Geometry group at ETH Zurich under the supervision of Marc Pollefeys; his thesis, On Learning and Geometry for Visual Localization and Mapping, was awarded the ETH Silver Medal and the DAGM MVTec Dissertation Award at GCPR 2025. He is best known for the learned feature matchers SuperGlue (CVPR 2020) and LightGlue (ICCV 2023), and for hloc, a widely used open-source toolbox for visual localization and structure-from-motion. His further work includes Pixel-Perfect Structure-from-Motion (ICCV 2021), the LaMAR benchmark for AR localization and mapping, and OrienterNet, SNAP and GeoCalib.
Andrea Vedaldi is Professor of Computer Vision and Machine Learning at the University of Oxford, where he co-leads the Visual Geometry Group (VGG) in the Department of Engineering Science. His research develops methods that understand the content of images and videos automatically, with little to no manual supervision, both in terms of semantics and of 3D geometry. He was elected a Fellow of the Royal Academy of Engineering in 2025 and is one of the inaugural recipients of the Royal Society Faraday Discovery Fellowship, which supports a long-term programme on spatial artificial intelligence. His recent work includes the VGGT family of feed-forward 3D reconstruction models, whose quality scales with data and model size; VGGT-Ω was named a best paper finalist at CVPR 2026.
Patrick Wenzel is an AI Research Engineer at Helsing in Munich. He obtained his PhD in computer science from the Technical University of Munich, advised by Daniel Cremers, with a focus on robust visual localization and mapping under long-term appearance change. He is a co-author of the 4Seasons dataset for multi-weather SLAM and long-term visual localization in automotive settings, and has worked on cross-season place recognition and learned front-ends for direct visual odometry, previously as a researcher at Artisense. His current interests lie in bringing learned perception and mapping systems into reliable real-world deployment.
Christian Wolf is a Principal Scientist at Naver Labs Europe, where he leads the Spatial AI team. From 2005 to 2021 he was an associate professor (Maître de Conférences, HDR) at INSA Lyon and the CNRS laboratory LIRIS, where he held the ANR/Naver/INSA chair in artificial intelligence “REMEMBER — Learning Reasoning, Memory and Behavior”. He received his MSc from TU Vienna in 2000, his PhD from INSA Lyon in 2003, and his habilitation in 2012. His research focuses on AI for robotics, in particular machine learning and embodied computer vision, the large-scale learning of high-level reasoning from visual observations, and the connections between machine learning and control. He is an ELLIS member and served as an associate editor of IEEE TPAMI from 2019 to 2025.
Organizers
Luca Carlone
Associate Professor
Massachusetts Institute of Technology
Qing Cheng
PhD Student
Technical University of Munich
Daniel Cremers
Professor
Technical University of Munich
Dima Damen
Professor
University of Bristol
Frank Dellaert
Professor
Georgia Institute of Technology
Ayoung Kim
Professor
Seoul National University
Patrick Wenzel
AI Research Engineer
Helsing
Niclas Zeller
Professor
Karlsruhe University of Applied Sciences
Related Workshops
Other ECCV 2026 workshops on closely related topics:
- 3rd Neural SLAM Workshop (NeuSLAM) — September 8, afternoon
- Privacy-Preserving Visual Localization and Mapping — September 8, afternoon
- Structure-from-Motion in the Age of Deep Learning (SfM-ADL) — September 9, morning
ViLMa takes place on the morning of September 8, so none of these clash with our programme.