Résumé
Global visual geolocation consists in predicting where an image was captured anywhere on Earth. Since not all images can be localized with the same precision, this task inherently involves a degree of ambiguity. However, existing approaches are deterministic and overlook this aspect. In this paper, we propose the first generative approach for visual geolocation based on diffusion and flow matching, and an extension to Riemannian flow matching, where the denoising process operates directly on the Earth's surface. Our model achieves state-of-the-art performance on three visual geolocation benchmarks: OpenStreetView-5M, YFCC100M, and iNat21. In addition, we introduce the task of probabilistic visual geolocation, where the model predicts a probability distribution over all possible locations instead of a single point. We implement new metrics and baselines for this task, demonstrating the advantages of our generative approach. Codes and models are available here.
| langue originale | Anglais |
|---|---|
| Pages (de - à) | 23016-23026 |
| Nombre de pages | 11 |
| journal | Proceedings of the IEEE Computer Society Conference on Computer Vision and Pattern Recognition |
| Les DOIs | |
| état | Publié - 1 janv. 2025 |
| Evénement | 2025 IEEE/CVF Conference on Computer Vision and Pattern Recognition, CVPR 2025 - Nashville, États-Unis Durée: 11 juin 2025 → 15 juin 2025 |
Empreinte digitale
Examiner les sujets de recherche de « Around the World in 80 Timesteps: A Generative Approach to Global Visual Geolocation ». Ensemble, ils forment une empreinte digitale unique.Contient cette citation
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver