Visual Odometry
Visual odometry (VO) computes the change in a vehicle’s six degree of freedom pose by tracking terrain features between two stereo image pairs taken before and after a short motion. On a planetary rover it is the only sensor able to measure positional slip, which wheel encoders cannot observe and which double integration of accelerometer output is too noisy to recover [1][3].
What it computes
Section titled “What it computes”Input is two stereo pairs, plus an initial motion guess from wheel odometry and the inertial measurement unit. Output is a corrected translation and rotation, or a refusal. The MER implementation applied only the position component of the update, because the Litton LN-200 IMU drifted less than 3 degrees per hour of operation and already held attitude well [1]. Curiosity kept the same configuration through its first seven years [3].
The MER and MSL algorithm is a 3D to 3D maximum likelihood estimator, in six steps [1]. A Forstner or Harris corner operator selects features on a grid over the left image, one per cell, enforcing a minimum feature separation. Each feature is then stereo matched along the epipolar line using pseudo-normalized correlation, with a biquadratic fit to a 3x3 neighborhood of correlation scores for subpixel position, and features whose left and right rays miss each other by more than a gap threshold are discarded [1]. Per-feature 3D covariance is computed from the curvature of that biquadratic fit, so a feature in low-texture terrain is weighted down relative to one in high texture. Features are projected into the second pair using the wheel odometry guess, re-matched by correlation search, and screened by a rigidity test on relative 3D distances that rejects gross outliers [1]. A weighted least squares pose solution with a closed form given by singular value decomposition is embedded in RANSAC, six features per sample, accepting a feature when its reprojection error is under 0.5 pixels. The RANSAC inlier set feeds a maximum likelihood estimate that uses the full 3x3 feature covariances, iterated until the rotation update falls below 6e-6 radians [1].
Convergence typically took 6.4 +/- 1.7 iterations on Spirit and 8.4 +/- 5.2 on Opportunity, tracking 73.4 +/- 29.3 and 87.4 +/- 34.1 features per step respectively [1].
Assumptions inherited from the founding algorithms
Section titled “Assumptions inherited from the founding algorithms”The outlier rejection in step 5 is RANSAC, whose trial count is k = log(1-z)/log(1-w^n) for confidence z of drawing at least one all-inlier sample of size n, with w the probability that any one datum is an inlier [9]. Both inputs are assumed rather than measured: w has to be known in advance, and the derivation requires samples drawn independently and uniformly, which stereo feature matches do not satisfy because mismatches cluster on repeated texture and on the low-texture patches that produce most flight failures. The standard deviation of the trial count is approximately equal to its expectation, so the original recommendation is to budget two to three times the expected number of trials, and the consensus threshold rule that five points beyond the minimal sample gives better than 95 percent confidence rests on an assumed probability below 0.5 that a point agrees with a wrong model, a quantity stated to be not generally determinable [9]. A fixed six-feature sample and a fixed 0.5 pixel reprojection tolerance are the flight expression of those assumptions.
Differential registration is the other half of the inheritance. Estimating a displacement by linearizing image intensity about the current estimate rests on a first-order Taylor expansion plus brightness constancy, and its convergence is proven only for a pure sinusoid, where it holds for initial misregistration up to half a wavelength; for a real image the stated requirement is only that the initial estimate locate the target to within about its own size [10]. That requirement is what the 60 percent overlap rule, the 75 cm step limit and the wheel odometry seed exist to satisfy. Smoothing widens the capture range and destroys any feature smaller than the smoothing window, a tradeoff named without being quantified, and the reported cost advantage, O(M^2 log N) against O(M^2 N^2) for exhaustive search over an M by M disparity range on an N by N image, holds only when the iteration converges [10].
Treating image motion as scene motion requires a flat surface, uniform incident illumination and smoothly varying reflectance with no spatial discontinuities [11]. Both of the named failures of that identification occur on Mars: a stationary scene under moving illumination changes without moving, which is the mechanism behind the residual left when the rover and mast are held still and only the micro-shadows near pebbles track the Sun [3], and a texture-free surface that moves without changing its image, which is feature starvation in sand. The 1 percent accuracy figure often quoted for the regularized flow solution is the average of the flow field over a whole synthetic 32 by 32 image in pure translation at about 1 percent added noise; the per-vector error in the same test is about 10 percent after 32 iterations on a single time step and about 7 percent at one iteration per time step, and the largest errors fall on occluding boundaries and where the brightness gradient is small [11].
Cost on flight hardware
Section titled “Cost on flight hardware”The processing cost, not the algorithm, set the drive rate on every mission before Mars 2020.
| Sojourner | Spirit and Opportunity | Curiosity | Perseverance | |
|---|---|---|---|---|
| Landing year | 1997 | 2004 | 2012 | 2021 |
| CPU | 80C85 | BAE RAD6000 | BAE RAD750 | BAE RAD750 x2 plus Xilinx Virtex5QV FPGA |
| Clock | 2 MHz | 20 MHz | 133 MHz | 133 MHz x2, FPGA at 22M disparities/s |
| RAM | 0.56 MB | 128 MB | 128 + 512 MB | 128 MB x2 plus 512 MB x2 |
| Non-volatile storage | 0.17 MB | 256 MB flash | 4096 MB flash | 4096 + 3072 MB flash |
| Image pairs per step | 1 | 1 to 2 | 4 | 1 |
| Stereo pixels per step | 20 | 10,000 to 50,000 | 40,000 to 200,000 | 240,000 to 1,200,000 |
| Autonomous navigation pause per step | not reported | about 120 s | about 120 s | typically 0 if not steering |
Source: [4], Table 1.
On MER at most 75 percent of the 20 MHz RAD6000 was available to autonomy software, with telemetry processing sometimes taking more [2]. Dynamic allocation was discouraged by the coding standards, leaving dedicated pools of 4 MB, 9 MB and up to ten further 2 MB blocks, and using those 2 MB blocks reduced the memory left for image processing [2]. A single VO tracking step averaged nearly three minutes of computation on the 20 MHz RAD6000, and took up to three minutes for one tracking step under the nominal mission software [1]. Driving with VO running continuously capped the maximum drive rate at around 10 m/h.
Curiosity’s RAD750 at 133 MHz with 512 MB of RAM cut the figure to 47 s per step on average, covering image acquisition, transfer to the CPU, downsampling, motion estimation and data product writing, a figure that includes the camera design, the camera interface board design and flash memory operation as well as the estimator [3]. VO Full drives averaged about 46 m/h against over 100 m/h for directed drives, with the CPU idle for much of each step because the rover stopped to think [5].
VO Thinking While Driving, developed from 2015 and validated over a three year campaign, starts the next drive step immediately after image capture so the motion estimate is computed while the wheels turn [5]. Measured on the flight-like engineering model, nine runs gave 51.5 m/h average without the capability and 77.5 m/h with it, a factor of about 1.5, with individual runs improving from 52.6 to 80.9 m/h on the same day [5]. The cost is localization latency: a large slip is not detected until two 1 m steps have completed rather than one.
Perseverance carries a second RAD750 board, the Vision Compute Element, whose Virtex-5 FPGA runs the stereo correlation and VO image processing [4][6]. It processes about six times as many pixels per step as Curiosity in less time, and for the first time the autonomous drive rate is limited by the wheel drive motor rotation rate rather than by sensing or computing [4]. Maximum drive speed remains 4.2 cm/s, fixed by gear ratios chosen for climbing rather than speed.
Measured performance
Section titled “Measured performance”| Vehicle | Period | Attempts | Convergence | Accuracy | Source |
|---|---|---|---|---|---|
| Spirit | to sol 414 | 609 evaluation steps | 590, 97 percent | slips to 125 percent measured, tilts 2 to 30 deg | [1] |
| Opportunity | to sol 394 | 875 evaluation steps | 828, 95 percent | tilts 0.8 to 31 deg, mean 18.0 +/- 4.6 deg | [1] |
| Curiosity | to sol 2488 | 20,682 attempts while driving | 20,588, 99.55 percent | 1.5 +/- 0.9 mm 3D position error over 141 static repeats | [3] |
| Perseverance | to sol 1710 | not reported as a count | 99.83 percent | over 99 percent of driving used VO | [6] |
The Curiosity accuracy figure comes from a sol 571 test in which the rover and remote sensing mast were held stationary and VO was run 141 times over 90.5 minutes against a single seed image, from 11:40:45 to 13:08:49 local mean solar time [3]. Error grew roughly linearly with the interval between seed and update image, and the residual is attributed to micro-shadows near pebbles moving with the Sun while the seed image did not [3]. During driving the seed and update images are about one minute apart, and the relocalization precision is better than 2 percent of distance driven [5].
On a 2.45 m rock-laden test course driven in 35 cm steps on the MER engineering model, with slip up to 85 percent, IMU and wheel odometry alone exceeded the 10 percent position error design goal after 1.4 m, while VO error stayed under 1 percent of distance traveled [1]. On Opportunity sols 188 to 191 a real 19 m uphill and cross-slope drive was underestimated by wheel odometry by 1.6 m, and the two final position estimates differed by nearly 5 m.
The comparable terrestrial algorithm of the same period, using inlier detection over dense disparity images rather than RANSAC over sparse features, ran in about 20 ms on a 512x384 image and held position error under 1 m after 4000 frames and 400 m of travel, 0.25 percent of distance [8]. The flight figures are three to four orders of magnitude slower for the same class of computation.
Failure modes
Section titled “Failure modes”Non-convergence has three flight-observed causes: too little image overlap after a large motion, too few trackable features in the terrain, and internal sanity checks rejecting an otherwise correct estimate [1][3].
Overlap is an operational constraint. MER split drives into steps of no more than 75 cm in a straight line or arc, and no more than 18 degrees of heading change per turn-in-place step, to hold at least 60 percent overlap between adjacent 256x256 NavCam pairs taken from 1.5 m above the ground [1]. A commanded 40 degree turn in place left too little overlap and forced re-initialization. Curiosity uses the same 60 percent overlap target with a 1 m nominal step, cameras at 1.96 m on the pointable mast, typically pitched 20 to 30 degrees below horizontal [3].
Feature starvation dominates the flight record. Of Curiosity’s 94 VO failures in seven years, 66 were failures to converge, most of them with the cameras pointed at all-sandy terrain [3]. The MER feature detector was tuned for feature-rich natural terrain and regularly failed in piles of sand; the workaround was that such terrain was pliable enough that the rover’s own tracks supplied the features [1]. On Perseverance, three drives over 75 m of visually bland terrain at Lookout Hill on sols 1339, 1342 and 1347 produced 34 VO failures, 42 percent of all VO failures in the mission to sol 1710, and faulted the drives on the first two sols [6].
Curiosity’s other 28 failures split into 12 step truncation reimaging failures caused by a flight software defect that suppressed re-imaging when the rover autonomously commanded a turn sharp enough to destroy overlap, 10 sequencing failures where planners knowingly sent commands outside recommended bounds, 3 strategy failures where an unrelated fault left the cameras pointing inconsistently, and 3 IMU parameter failures on sols 122 to 124 [3]. The last of these were false rejections: the built-in check requiring VO and IMU attitude to agree within 1 degree in roll, pitch and yaw tripped on pitch differences of -1.391, -1.072 and -2.195 degrees, and the subsequent investigation concluded the VO update was correct and the IMU parameter settings were dropping attitude updates under high acceleration.
False positives have occurred. On Opportunity sols 137 and 141 VO converged to unreasonable position updates because the minimum feature separation parameter was set too small, letting the detector cluster features in one small planar feature-rich patch [1]. The response was a set of optional constraints added in the February 2005 flight software release, letting rover planners bound the update magnitude, its X and Y world frame components, the roll, pitch and yaw changes, the angle from downslope above a given tilt, and the number of tolerated non-convergences.
Consequences for driving
Section titled “Consequences for driving”VO turns an unobservable quantity into a fault threshold. Curiosity computes a wheel slip fraction, the summed linear distances of all VO-corrected wheel positions from their no-slip positions divided by summed wheel path length, and a rover slip fraction, the Euclidean distance between the no-slip and VO-corrected rover origin divided by no-slip path length, the latter undefined during turns in place [3]. Over the mission to sol 2488 average rover slip was 6.24 percent, average wheel slip 8.44 percent, and the maximum of both was 98.70 percent on sol 2087 [3]. Five drives were stopped by VO failures and the worst single sol, 2434, saw 11 failures against a configured limit of 10, stopping the drive 2.82 m short.
Opportunity’s embedding in the Purgatory ripple established the operational pattern. After 50 m of blind driving that produced about 2 m of track, VO measured slip rates of 98.9 to 99.5 percent over sols 463 to 483, in one case about 1 mm of motion against 2 m commanded [1]. The extrication procedure ran VO after each commanded 2 m and halted all driving when VO either measured non-trivial motion or failed to converge three times. Slip Check then became standard: blind segments capped at 5 m, followed by a 20 cm drive step with VO to confirm the vehicle can still move, which bounds how deeply the wheels can bury before anyone notices [1]. Perseverance inverted the same logic on sol 1347, disabling the standard VO failure fault protection and requiring VO to succeed only once every 20 m, justified by the mechanical specification that the rover can extract itself after 20 m of motion without forward progress [6].
Position uncertainty growth is the budget VO feeds. Perseverance is configured to assume 5 percent of distance traveled when VO converges, against a demonstrated potential of 2 percent per 100 m, and 50 percent of distance traveled for any motion where VO failed to converge [7]. That weighted sum determines how far human-specified keep-out zones must be grown on the orbital map, and it is what onboard global localization against orbital imagery exists to reset [7]. The localization itself is too expensive for the Rover Compute Element and was moved to the Ingenuity helicopter base station, where four Snapdragon cores at 2.36 GHz run it in about 32 s against an order of magnitude longer on the 133 MHz RAD750.
Usage share follows from the cost. MER used onboard terrain assessment for 1354 of Spirit’s 4798 m and 1379 of Opportunity’s 5947 m as of 15 August 2005, about 25 percent for both [2]. Curiosity drove over 88 percent of its distance under some form of VO, mostly VO Full [3]. Perseverance had driven 31,857 m in AVOID_ALL mode at an effective 92.67 m/h and 9910 m unguarded at 82.45 m/h through sol 1710, 75.28 percent of odometry under full autonomy, with a single-sol record of 403 m autonomous out of 411.7 m on sol 1540 [6]. During the 31 sol Rapid Traverse Campaign of March 2022, 94.8 percent of 5063.4 m was planned onboard at an average 79.7 m/h over an average 2.98 h of driving per sol [4].
References
- Maimone, M., Cheng, Y. and Matthies, L. (2007). Two Years of Visual Odometry on the Mars Exploration Rovers. Journal of Field Robotics, 3. Source
BibTeX
@article{maimone2007two, title = {Two Years of Visual Odometry on the Mars Exploration Rovers}, author = {Maimone, Mark and Cheng, Yang and Matthies, Larry}, institution = {NASA Jet Propulsion Laboratory}, year = {2007}, journal = {Journal of Field Robotics}, doi = {10.1002/rob.20184}, volume = {24}, pages = {169--186}, number = {3}, url = {https://www-robotics.jpl.nasa.gov/media/documents/rob-06-0081.R4.pdf} } - Maimone, M. W., Leger, P. C. and Biesiadecki, J. J. (2007). Overview of the Mars Exploration Rovers' Autonomous Mobility and Vision Capabilities. Source
BibTeX
@inproceedings{maimone2007overview, title = {Overview of the Mars Exploration Rovers' Autonomous Mobility and Vision Capabilities}, author = {Maimone, Mark W. and Leger, P. Chris and Biesiadecki, Jeffrey J.}, booktitle = {IEEE International Conference on Robotics and Automation, Space Robotics Workshop}, address = {Rome, Italy}, year = {2007}, url = {https://www-robotics.jpl.nasa.gov/media/documents/mer_autonomy_icra_2007.pdf} } - Rankin, A., Maimone, M., Biesiadecki, J., Patel, N., Levine, D. and Toupet, O. (2021). Mars Curiosity Rover Mobility Trends During the First Seven Years. Journal of Field Robotics, 5. Source
BibTeX
@article{rankin2021mars, title = {Mars Curiosity Rover Mobility Trends During the First Seven Years}, author = {Rankin, Arturo and Maimone, Mark and Biesiadecki, Jeffrey and Patel, Nikunj and Levine, Dan and Toupet, Olivier}, year = {2021}, journal = {Journal of Field Robotics}, volume = {38}, number = {5}, pages = {759--800}, doi = {10.1002/rob.22011}, url = {https://www-robotics.jpl.nasa.gov/media/documents/ROB-20-0040_R3.pdf} } - Rankin, A., Del Sesto, T., Hwang, P., Justice, H., Maimone, M., Verma, V. and Graser, E. (2023). Perseverance Rapid Traverse Campaign. Source
BibTeX
@inproceedings{rankin2023perseverance, title = {Perseverance Rapid Traverse Campaign}, author = {Rankin, Arturo and Del Sesto, Tyler and Hwang, Pauline and Justice, Heather and Maimone, Mark and Verma, Vandi and Graser, Evan}, booktitle = {2023 IEEE Aerospace Conference}, address = {Big Sky, Montana}, year = {2023}, url = {https://robotics.jpl.nasa.gov/media/documents/2023-rapid-traverse.pdf}, doi = {10.1109/aero55745.2023.10115835}, pages = {1-16} } - Rankin, A., Holloway, A., Sabel, A., Patel, N. and Maimone, M. W. (2022). Visual Odometry Thinking While Driving for the Curiosity Mars Rover's Three-Year Test Campaign: Impact of Evolving Constraints on Verification and Validation. NASA, 20230005759. Source
BibTeX
@inproceedings{rankin2022visual, title = {Visual Odometry Thinking While Driving for the Curiosity Mars Rover's Three-Year Test Campaign: Impact of Evolving Constraints on Verification and Validation}, author = {Rankin, Arturo and Holloway, Alexandra and Sabel, Anna and Patel, Nikunj and Maimone, Mark W.}, year = {2022}, institution = {NASA}, number = {20230005759}, url = {https://ntrs.nasa.gov/citations/20230005759}, booktitle = {2022 IEEE Aerospace Conference (AERO)}, doi = {10.1109/aero53065.2022.9843487}, pages = {1-10} } - Maimone, M., Verma, V., Rankin, A., Kaplan, K., Carsten, J., Schaler, E., Boroson, E., Graser, E., Srinivasan, T., Nash, J. and Chiu, D. (2026). Roving on the Edge: Robotic Operations Power Perseverance's Ascent of Jezero Crater Rim. Source
BibTeX
@inproceedings{maimone2026roving, title = {Roving on the Edge: Robotic Operations Power Perseverance's Ascent of Jezero Crater Rim}, author = {Maimone, Mark and Verma, Vandi and Rankin, Arturo and Kaplan, Kyle and Carsten, Joseph and Schaler, Ethan and Boroson, Elizabeth and Graser, Evan and Srinivasan, Thirupathi and Nash, Jeremy and Chiu, Darwin}, booktitle = {AAS Guidance, Navigation and Control Conference}, year = {2026}, url = {https://www-robotics.jpl.nasa.gov/media/documents/2026_RO_AAS_final.pdf} } - Verma, V., Nash, J., Saldyt, L., Dwight, Q., Wang, H., Myint, S., Biesiadecki, J., Maimone, M., Tumbar, A., Ansar, A., Kubiak, G. and Hogg, R. (2024). Enabling Long and Precise Drives for the Perseverance Mars Rover via Onboard Global Localization. Source
BibTeX
@inproceedings{verma2024enabling, title = {Enabling Long and Precise Drives for the Perseverance Mars Rover via Onboard Global Localization}, author = {Verma, Vandi and Nash, Jeremy and Saldyt, Lucas and Dwight, Quintin and Wang, Haoda and Myint, Steven and Biesiadecki, Jeffrey and Maimone, Mark and Tumbar, Andrei and Ansar, Adnan and Kubiak, Gerik and Hogg, Robert}, booktitle = {IEEE Aerospace Conference}, address = {Big Sky, Montana}, year = {2024}, url = {https://www-robotics.jpl.nasa.gov/media/documents/2024_Global_Localization_IEEE_Aero.pdf} } - Howard, A. (2008). Real-Time Stereo Visual Odometry for Autonomous Ground Vehicles. Source
BibTeX
@inproceedings{howard2008real, title = {Real-Time Stereo Visual Odometry for Autonomous Ground Vehicles}, author = {Howard, Andrew}, booktitle = {IEEE/RSJ International Conference on Intelligent Robots and Systems}, address = {Nice, France}, year = {2008}, url = {https://www-robotics.jpl.nasa.gov/media/documents/howard_iros08_visodom.pdf} } - Fischler, M. A. and Bolles, R. C. (1981). Random Sample Consensus: A Paradigm for Model Fitting with Applications to Image Analysis and Automated Cartography. Communications of the ACM, 6. Source
BibTeX
@article{fischler1981random, title = {Random Sample Consensus: A Paradigm for Model Fitting with Applications to Image Analysis and Automated Cartography}, author = {Fischler, Martin A. and Bolles, Robert C.}, journal = {Communications of the ACM}, volume = {24}, number = {6}, pages = {381--395}, year = {1981}, doi = {10.1145/358669.358692}, url = {https://www.sri.com/wp-content/uploads/2021/12/ransac-publication.pdf} } - Lucas, B. D. and Kanade, T. (1981). An Iterative Image Registration Technique with an Application to Stereo Vision. Source
BibTeX
@inproceedings{lucas1981iterative, title = {An Iterative Image Registration Technique with an Application to Stereo Vision}, author = {Lucas, Bruce D. and Kanade, Takeo}, booktitle = {Proceedings of the DARPA Image Understanding Workshop}, pages = {121--130}, year = {1981}, url = {https://www.ri.cmu.edu/pub_files/pub3/lucas_bruce_d_1981_2/lucas_bruce_d_1981_2.pdf} } - Horn, B. K. P. and Schunck, B. G. (1981). Determining Optical Flow. Artificial Intelligence, 1--3. Source
BibTeX
@article{horn1981determining, title = {Determining Optical Flow}, author = {Horn, Berthold K. P. and Schunck, Brian G.}, journal = {Artificial Intelligence}, volume = {17}, number = {1--3}, pages = {185--203}, year = {1981}, doi = {10.1016/0004-3702(81)90024-2}, url = {https://people.csail.mit.edu/bkph/papers/Optical_Flow_OPT.pdf} }