• About
  • Competitions
  • Lunar X-view
  • Register
  • Sign In
    Forgot password?
    GitHub
    By logging in with Github you accept our Privacy Policy
Lunar X-view
  • Home
  • Challenge
  • Data
  • Scoring
  • Rules
  • Submission Rules
  • Discussion
  • Resources
Timeline

Nov. 1, 2026, midnight UTC

March 31, 2027, midnight UTC

Challenge

In a nutshell

Estimate the position of a panorama picture taken from the Moon surface (such as the one above), from an on-board aerial map of (a rather large) area.

Background

Lunar rovers have an awkward navigation problem: they are good at knowing how they have moved, but not always at knowing exactly where they are. Wheel odometry, inertial sensors, visual odometry, and local maps can track motion accurately over short distances; over longer traverses, however, small errors accumulate. Eventually, the rover’s internal map and the global lunar map start to disagree: never good when the nearest road sign is several hundred thousand kilometres away.

A natural way to reset this accumulated error is to compare what the rover sees on the ground with what an orbiter has already mapped from above. This is the core idea behind cross-view geolocalisation: given an image from one viewpoint, identify the corresponding location in imagery acquired from a very different viewpoint.

On Earth, cross-view geolocalisation has become a major computer-vision problem: match a street-level image against aerial or satellite imagery, and infer where it was taken [1–3]. It has obvious applications in navigation where GNSS is unavailable, unreliable, or simply not enough. Lunar X-View takes this familiar “find this place from another viewpoint” problem and removes many of the convenient Earth features: roads, buildings, signs, vegetation, and ... coffee shops (they are prominent landmarks in many places).

Test your human eye.
Is the panorama above from this map?

The Moon replaces them with craters, boulders, ridges, scarps, regolith texture, and shadows that can be spectacularly unhelpful. A terrain patch seen from a rover can look radically different from the same patch seen from orbit: the viewpoints differ by orders of magnitude, the scales and perspectives do not match, most importantly the illumination geometry changes during the lunar day, and local topography creates occlusions and long shadows. In some cases, the landscape is also visually repetitive enough that several locations may look plausible.

This is not merely an academic exercise. Planetary missions need global localisation to support long autonomous traverses, global-map alignment, science-targeting, hazard assessment, and coordination among surface assets. On Mars, Nash et al. demonstrated this operational value proposing Censible,  a system that matches Perseverance rover panoramas with orbital maps to correct accumulated localisation error [4]. Their approach achieved a mean error of 0.36 m on 264 real Mars panoramas and was designed for onboard, resource-constrained operation. Lunar localisation has also been investigated using rover-orbiter imagery and crater-based landmarks [5, 6].

Censible demonstrates that rover-to-orbiter matching can provide accurate onboard localisation when mission-specific stereo panoramas, high-resolution orbital imagery, elevation maps, an orientation estimate, and a bounded search region are available [4]. Lunar X-View addresses a complementary and more open-ended problem. It evaluates whether a method can retrieve the correct location over a substantially larger candidate area, remain reliable when viewpoint and illumination conditions differ from those encountered during training, and most notably from those available from the aerial map and operate using appearance information alone, without orbital or rover-derived elevation data. The objective is therefore not to replace mission-specific geometric localisation systems, but to assess the quality, robustness and transferability of cross-view methods when strong operational priors or auxiliary terrain information are unavailable. 

The Lunar X-View roots are in the broader Earth-based cross-view-localisation community; its playground is the Moon, where the visual problem is harsher, the navigation stakes are higher, and there is no option to ask a pedestrian for directions.

Problem description

Our Lunar rover is lost on the Moon with no tracking, elevation data or path history available. The only hope of recovering its position is it's camera on a periscope that can take a 360° panoramic image 2.5m above ground. The lunar night is approaching and the shadows of the craters grow larger and larger: once bright areas on its map are now completely blacked out by shadows. The potential area is vast, approximately 750km² provided as a single large map displaying a bird's eye view camera with a 5m/pixel resolution, pointing straight down towards the surface.

All images, panoramas and maps have been generated using the Planet and Asteroid Natural Scene Generation Utility (PANGU) developed by the University of Dundee’s Space Technology Centre, an established software for planetary science simulations. We choose the Hapke BRDF model (Bidirectional Reflectance Distribution Function) mixed with a small Lambertian diffusion term to provide realistic reflectances from the porous and rough regolith surface. This allows us to model subtle, optical phenoma such as the opposition effect introducing physically plausible sources of confusion to the images.

Small-scale features such as ridges and Moon rocks are scattered following realistic distributions and correlations with regards to crater positions. Furthermore, we apply a high-resolution texture extracted from public footage of the Chinese Yutu Rover and several Noise terms to increase the realism of the surface and imaging systems.

Although our data does not contain real moon footage, it was generated with high-resolution 5m/px digital elevation maps (DEMs) from actual lunar regions at its foundation. While DEMs play a significant role in localization and navigation, they also impose challenges and limitations for processing pipelines due to their size compared to camera images. Moreso, future missions to planetary bodies might not have high-resolution DEMs available at all. Thus, in alignment to terrestrial applications of cross-view localization, Lunar X-view does not provide any DEM and is intentionally designed as a from pixels only computer vision challenge.

While reconstructing DEMs could be one of many potential strategies to tackle this challenges, we discourage competitors to identify the foundational DEM of our Lunar regions, which is why we heavily modified it, retaining the general geomorphologies and features of the moon but perturbing in a way that it does no longer relates to a specific region of the Moon. Instead, two different pseudo-regions are provided for you, one as a reference with location labels and one with only a few examples, leaving the rest to your skills. Our dataset section contains more specifics about the contents of our dataset.

Your method to cross-reference the panoramic images to the corresponding maps can be anything: all we ask of you is to tell us at which position (x,y) you think that the panoramas were taken (see submission rules for details). The closer your guesses are to the actual locations, the better your score will be. Will you be able to reach absolute zero by pinpointing all 10 000 panoramas exactly?

Eye candy

This video (aerial to panoramic view) was rendered using our low-res pipeline, to show a flight from an orbit view to panorama conditions. 


References

[1] A. Durgam, S. Paheding, V. Dhiman, and V. K. Devabhaktuni, “Cross-view geo-localization: A survey,” IEEE Access, vol. 12, pp. 168614–168642, 2024. Also available as arXiv:2406.09722.

[2] S. Workman, R. Souvenir, and N. Jacobs, “Wide-area image geolocalization with aerial reference imagery,” in Proceedings of the IEEE International Conference on Computer Vision (ICCV), Santiago, Chile, 2015, pp. 3961–3969.

[3] L. Liu and H. Li, “Lending orientation to neural networks for cross-view geo-localization,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Long Beach, CA, USA, 2019, pp. 5624–5633.

[4] J. Nash, Q. Dwight, L. Saldyt, H. Wang, S. Myint, A. Ansar, and V. Verma, “Censible: A robust and practical global localization framework for planetary surface missions,” in Proceedings of the IEEE International Conference on Robotics and Automation (ICRA), Yokohama, Japan, 2024, pp. 8642–8648.

[5] X. Zhao, L. Cui, X. Wei, C. Liu, J. Yin, H. Zhang, and Y. Gao, “Lunar rover cross-view localization through integration of rover and orbital images,” IEEE Transactions on Geoscience and Remote Sensing, vol. 62, 2024.

[6] S. Silvestrini, M. Piccinin, G. Zanotti, A. Brandonisio, I. Bloise, L. Feruglio, P. Lunghi, M. Lavagna, and M. Varile, “Optical navigation for lunar landing based on convolutional neural network crater detector,” Aerospace Science and Technology, vol. 123, Art. no. 107503, 2022.

 

 

Created by the Advanced Concepts Team, Copyright © European Space Agency 2021