Skip to content

Flight Software Verification

Flight software verification is a ladder of test venues ordered by fidelity and inversely by throughput. The bottom rung is a workstation running the flight software against simulated hardware; the top is the flight vehicle itself. The engineering question at every rung is which defect classes it can find and which it structurally cannot.

MSL’s internal testing described four stages: unit tests, the Work Station Test Set, System Integrated Test, and Assembly Test and Launch Operations on real spacecraft hardware, with cost rising and availability falling at each step [4]. The Work Station Test Set runs the flight software under a VxWorks simulator on a Linux workstation and is described as cheap relative to every other method.

VenueFidelitySpeedMission
Work Station Test Set (WSTS)flight software against modeled avionics, no real hardware2 to 10 times real time depending on scenario complexity [3]MSL
WSTS, Mars 2020 buildsame, enhanced6 times real time [5]Mars 2020
SSimflight software in the loop with rover mechanisms and terrainup to 1000 times faster than the flight compute element [1]MSL, Mars 2020
Mission System Testbed (MSTB)hardware in the loop with two Rover Compute Elements, no actuators [5]real timeMars 2020
Vehicle System Testbed (VSTB)near-duplicate rover [5]real timeMars 2020
Integrated Test Laboratory (ITL)system-level, mixed real and simulated units [6]real timeCassini

The workstation simulator earns its place on usage, not on fidelity. MSL’s WSTS was first deployed in October 2006 and had reached its thirtieth release by 2008, by which point its monthly usage peaked near 3500 hours and it averaged an order of magnitude more use than the project’s hardware-in-the-loop testbeds [3]. It models every hardware behavior visible to the flight software as finite state machines, integrating differential equations where the dynamics require it, and supports fault injection at either the functional level or by setting individual bits, along with pause, resume and single step.

Mars 2020 built the Surface System Development Environment for the same reason inverted: Rover Compute Elements and motor controller assemblies were scarce and shared across subsystems, and the physical test venues, running three shifts a day seven days a week, still could not absorb the demand [2]. The sampling and caching subsystem alone has 19 actuators and over 100 sensors, including strain gauges, resolvers, temperature sensors and contact switches. The environment presents one interface to the flight software over three interchangeable backends: an abstraction layer giving deterministic force and motion simulation, a terrain simulator with slip modeling, visual odometry, pose tracking and rendered camera imagery, and a motor server that drives real testbed hardware over TCP/IP. Substituting the backend rather than the test case is what lets the same test run in software and on hardware.

What simulation found that nothing else would have

Section titled “What simulation found that nothing else would have”

Two flight-software state interactions on MSL are the documented catches. In one, a MastCam focus mechanism lock was not properly monitored before a motion command, and a visual odometry command unconditionally reinforced that state even though it was not a drive command; SSim caught the improperly set focus lock that would have failed the entire drive, and SSim was extended to model MastCam as a result [1]. In the other, SSim was used to model ChemCam laser targeting in terrain-relative frames to check the beam against the rover body for self-intersection.

These are combinatorial state defects, and they are found in simulation because simulation is where the combinations can be run. MSL flight software carries more than 4200 commands with dozens of arguments each, 54,000 parameters and tens of thousands of state variables [1]. The arm at Rocknest alone executed 2404 moves, against 2303 for Opportunity over its entire Victoria Crater journey, and self-collision checking requires at least 3 mm of margin.

Hardware in the loop found a different class. Mars 2020 ran the development environment against brassboard hardware at three venues from 2017 and 2018, the docking testbed with robotic arm and dock hardware, a coring software venue with the corer, and a multi-station testbed with the caching assembly, which gave the flight software early runtime against real hardware at a point in the project when changing the software and the algorithms was still comparatively cheap [2]. Engineering model testing of the arm ran August 2018 to January 2019 covering requirements, preloading, docking and fault protection, engineering model caching assembly testing followed from December 2018, and the flight model caching assembly was run in a Mars-like thermal vacuum chamber in late 2019 and early 2020.

What the defect record says the testing misses

Section titled “What the defect record says the testing misses”

The one quantitative study of the question analyzed 199 critical post-launch anomalies across seven spacecraft, Galileo, Mars Global Surveyor, Cassini and Huygens, Deep Space 1, Mars Climate Orbiter, Mars Polar Lander and Stardust, using orthogonal defect classification: when the defect occurred, the trigger condition that surfaced it, the high-level entity fixed, and the actual fix [7]. Two results were unexpected to the authors. Outdated or missing procedures were involved in one fifth of the critical incidents, so the fix was to documentation rather than to code. And requirements changes were rarely because a previous requirement was wrong; they were new requirements for handling rare events or for compensating for hardware failures and degradation that had not been anticipated.

The same study found more problem reports classified as “nothing fixed” than expected. Those false positives, where test personnel believed the flight software was wrong when it was correct, turned out to forecast where operational personnel would later be confused: 311 developmental problem reports from MER avionics integration and test and assembly test and launch operations between April 2002 and January 2003 were analyzed on that basis, and the reports identified which features needed targeted training and documentation before launch [7]. The finding is about the operators rather than about the code.

What a fault injection campaign can and cannot report

Section titled “What a fault injection campaign can and cannot report”

The reference measurement on fault detection and recovery in a fault-tolerant flight computer is the FTMP engineering model campaign: 21,055 pin-level stuck faults injected on eight card types in one of ten line replaceable units, of which 17,418 were detected, every detected fault correctly identified and recovered by purging the module and taking a spare or degrading, and an average system reconfiguration time of 82 ms, with per-card averages from 46 ms for the system bus controller to 113 ms for PROM [9]. The floor under that recovery time is not the voters but the configuration control task running in the 3.125 Hz rate group, which is the same reason a spacecraft whose fault response is a scheduled task recovers in tens of milliseconds rather than microseconds.

The report refuses to turn the remaining 17.3 percent into a coverage figure, and the refusal is the transferable part: unused gates inside a package and unused signal pins are undetectable and also harmless, no classification separating them from real misses was done, so no coverage number is stated at all [9]. Two card types are missing from the data set because the only socketed CMOS sample was destroyed in handling, everything was injected in one unit of ten, and a stuck pin is a hardware failure model rather than an upset, a latch that nothing monitors, or a common-mode design fault.

Verifying software whose behavior is not enumerable

Section titled “Verifying software whose behavior is not enumerable”

The Mars 2020 onboard planner is the hardest case in the record because its output depends on the state the rover happens to be in. The verification campaign ran in spring 2022, after more than a year of Perseverance surface operations under the master and submaster scheme, and drew its scenarios from hundreds of sols of real planning [5]. Requirements were managed in DOORS Next Generation with a matrix linking each requirement to the verification activities providing its evidence, test setup and configuration were standardized in Jupyter notebooks, and the bulk of execution ran on WSTS at 6 times real time, with a lightweight planner simulator for cheaper cases, the Mission System Testbed for performance and thermal fidelity, and the Vehicle System Testbed for end-to-end runs. The planner entered primary operations on 5 October 2023 [5].

What the campaign explicitly could not cover is stated: the unbounded problem space of conditions the planner will meet in flight, which had to be constrained to a finite set of testable units; unknown deviations from predicted activity execution; and the behavioral uncertainty the flexible execution scheme introduces [5]. The anomalies closed during testing cluster around heating: managing onboard heating is difficult because temperatures must be monitored before heating starts, the thermal tables carry uncertainty, and heating can span a sleep cycle. The planner runs on a 133 MHz RAD750 rated at 266 MIPS, shared with the rest of the flight software [5].

Regression as the dominant cost after launch

Section titled “Regression as the dominant cost after launch”

Cassini’s methodology is the long-duration case: attitude control flight software implementation began in 1993, the launch load completed in 1997, and the prime mission ran from July 2004 [6]. Pre-launch testing was a ten-step sequence from requirements trace matrices through test plans, case design, test trace matrices, procedures, a readiness review, execution, anomaly reporting, analysis and reporting. Unit tests verified functions, modes and states and boundary conditions with interfaces stubbed; an integration baseline was configuration-controlled at the end of each phase; and acceptance testing verified the requirement test matrix with scenario and functional tests covering whatever the matrix did not.

The test environment was tiered like MSL’s but with real units mixed in. The flight software development simulator ran faster than real time on the same UNIX platform as the software; the subsystem test bench, and after launch the Cassini attitude control test station, ran real time with two flight computers, two inertial reference units, one real reaction wheel assembly plus simulated ones, and a simulated power subsystem; and the Integrated Test Laboratory ran system level in real time with two flight computers, two engineering flight computers, two stellar reference units, two Sun sensor assemblies, two main engine valve drivers, five engine gimbal actuators of which two were real, and two solid state recorders [6].

Each critical event got its own build and its own campaign: the A7 cruise build, then A8.6.7 for Saturn orbit insertion, A8.7.1 for probe release and relay, and A8.7.2 and its variants for the remainder of the prime mission [6]. Operations testing over the first two and a half years covered functionality and mission performance, stress testing to find the system’s limits, and review of uplink products and procedures. Full flight software uploads had become virtually non-existent by the prime mission, and the operating philosophy was conservative: do not change logic unless necessary. The same instinct governs the Mars 2020 phased delivery of onboard planning authority [5].

The Phoenix lander’s Payload Interoperability Testbed at the University of Arizona held a near-duplicate lander and arm and served a purpose no software simulator does: image comparison [8]. Before a sample delivery, the scoop was positioned over the selected instrument inlet port, a Robotic Arm Camera image was taken, and it was overlaid on an image of the same pose taken in the testbed, with positioning corrected from the difference. The testbed image is the reference the flight image is judged against, and no model substitutes for it because the discrepancy being looked for is exactly what the model does not know.

MethodCatchesStructurally misses
Workstation simulatorflight software logic, state machine interactions, fault responses to injected faults [3]data content of instruments it models only in format [3]; real-time behavior, since it deliberately runs off real time [3]
Flight software in the loop with mechanism modelscommand sequence errors, state interaction defects, self-collision [1]forces it averages from historical data rather than models [1]; instrument power, thermal and telecom behavior [1]
Hardware in the loop, brassboardinterface and timing behavior against real actuators early enough to change the algorithm [2]microsecond-scale bus responses the simulation coordinates on time slices [2]
Flight-like vehicleend-to-end behavior [5]the state combinations that only a real mission produces [5]
Post-launch defect analysisthat one fifth of critical incidents trace to procedures, not code [7]nothing, but only after the fact [7]

References

  1. Verma, V. and Leger, C. (2019). SSim: NASA Mars Rover Robotics Flight Software Simulation . IEEE Aerospace Conference. Source
    BibTeX
    @inproceedings{verma2019ssim,
      title = {SSim: NASA Mars Rover Robotics Flight Software Simulation},
      author = {Verma, Vandi and Leger, Chris},
      booktitle = {IEEE Aerospace Conference},
      pages = {1-11},
      address = {Big Sky, Montana},
      year = {2019},
      doi = {10.1109/aero.2019.8741862},
      abstract = {Each new Mars rover has pursued increasingly richer science while tolerating a wider variety of environmental conditions and hardware degradation over longer mission operation duration. Sojourner operated for 83 sols (Martian days), Spirit for 2208 sols, and Opportunity is at 5111 sols, and Curiosity operation is ongoing at 2208 sols. To handle this increase in capability, the complexity of onboard flight software has increased. MSL (also known as Curiosity), uses more flight software lines of code than all previous missions to Mars combined, including both successes and failures[1]. MSL has more than 4,200 commands with as many as dozens of arguments, 54,000 parameters, and tens of thousands of additional state variables. A single high-level command may perform hours of configurable robotic arm and sampling behavior. Incorrect usage can result in the loss of an activity or the loss of the mission. Surface Simulation (“SSim”) was developed to address the challenge of making full and effective use of many capabilities of MSL, while managing complexity and risk. SSim is software that performs rapid context sensitive simulation of flight software. NASA Mars missions are comprised of three phases: several months of Cruise, a brief but exciting Entry Descent and Landing (EDL), and a Surface mission that typically lasts as long as the hardware survives. SSim is meant for use during the surface phase when the mission fulfills its primary objectives. The focus of SSim on MSL was the robotic flight software, including rover mobility and navigation, robotic arm manipulation, and sample acquisition, processing, and delivery. It can execute behaviors in simulation a thousand times faster than they execute in real time on the flight compute element. SSim is used by rover drivers to develop and validate command sequences throughout the planning cycle. SSim has been used to plan all of the Curiosity robotic operations since landing and is expected to continue to be used for the remaining life of the rover. Due to the impact of SSim on MSL, the Mars 2020 mission plans to increase the scope of SSim during flight operations, simulating not only rover planner operations, but all surface operations, including the instrument, power, thermal and telecommunication behavior. SSim is part of the Rover Sequencing and Visualization (RSVP) suite of Rover Planning tools [2]. In the paper we provide an overview of SSim architecture, design, implementation, and usage on MSL, as well as an overview of plans for Mars 2020.}
    }
  2. Brooks, S., Litwin, T., Biesiadecki, J., Abcouwer, N., Del Sesto, T., McHenry, M., Myint, S., Twu, P. and Wai, D. (2022). Testing Mars 2020 Flight Software and Hardware in the Surface System Development Environment . IEEE Aerospace Conference. Source
    BibTeX
    @inproceedings{brooks2022testing,
      title = {Testing Mars 2020 Flight Software and Hardware in the Surface System Development Environment},
      author = {Brooks, Sawyer and Litwin, Todd and Biesiadecki, Jeffrey and Abcouwer, Neil and Del Sesto, Tyler and McHenry, Michael and Myint, Steven and Twu, Philip and Wai, Dennis},
      booktitle = {IEEE Aerospace Conference},
      pages = {1-13},
      address = {Big Sky, Montana},
      year = {2022},
      doi = {10.1109/aero53065.2022.9843794},
      abstract = {The Mars 2020 (M2020) Perseverance Rover is NASA's most advanced planetary rover mission to date. It includes a novel Sample Caching Subsystem (SCS) which will collect rock cores for possible future return to Earth, as well as an improved mobility system with enhanced autonomous navigation which will enable it to traverse faster and farther than prior rovers. The development of both systems required extensive flight software and flight hardware testing. To support this testing, we developed the Surface System Development Environment (SSDEV) and used it for a wide variety of testing. SSDEV is a bundled subset of M2020 Flight Software which runs on commercially available Linux computers and can be combined with multiple backend options for simulation and hardware control. The SSDEV architecture enabled our teams to perform much more testing of flight software and flight hardware than would have otherwise been possible. As a secondary benefit, the SSDEV-based test campaigns also helped our teams enter the operations phase of the mission with greater readiness of operations products and tools. In this paper, we summarize the motivation for SSDEV, provide an overview of the SSDEV architecture, list several examples of how SSDEV was used, and summarize lessons learned. SSDEV is not a substitute for integrated testing with flight-like avionics, but it enabled substantially more testing than would have otherwise been possible and also provided some unique benefits. We recommend architectures like SSDEV to future projects that need to perform extensive hardware and software testing using a limited set of flight-like avionics.}
    }
  3. Henriquez, D., Canham, T., Chang, J. T. and McMahon, E. (2008). Workstation-Based Avionics Simulator to Support Mars Science Laboratory Flight Software Development . AIAA SPACE Conference and Exposition. Source
    BibTeX
    @inproceedings{henriquez2008workstation,
      title = {Workstation-Based Avionics Simulator to Support Mars Science Laboratory Flight Software Development},
      author = {Henriquez, David and Canham, Timothy and Chang, Johnny T. and McMahon, Elihu},
      booktitle = {AIAA SPACE Conference and Exposition},
      year = {2008},
      doi = {10.2514/6.2008-6550},
      abstract = {The Mars Science Laboratory developed the WorkStation TestSet (WSTS) to support flight software development. The WSTS is the non-real-time flight avionics simulator that is designed to be completely software-based and run on a workstation class Linux PC. This provides flight software developers with their own virtual avionics testbed and allows device-level and functional software testing when hardware testbeds are either not yet available or have limited availability. The WSTS has successfully off-loaded many flight software development activities from the project testbeds. At the writing of this paper, the WSTS has averaged an order of magnitude more usage than the project's hardware testbeds.}
    }
  4. Jones, J. D. and Lam, D. (2011). Mars Science Laboratory Flight Software Internal Testing . NASA Jet Propulsion Laboratory, Undergraduate Student Research Program Final Report. Source
    BibTeX
    @techreport{jones2011mars,
      title = {Mars Science Laboratory Flight Software Internal Testing},
      author = {Jones, Justin D. and Lam, Danny},
      institution = {NASA Jet Propulsion Laboratory, Undergraduate Student Research Program Final Report},
      year = {2011},
      url = {https://dataverse.jpl.nasa.gov/dataset.xhtml?persistentId=hdl:2014/43488}
    }
  5. Parjan, S. and Gaines, D. (2024). In OBP We Trust: Verification and Validation of the M2020 On Board Planner Flight Software . IEEE Aerospace Conference. Source
    BibTeX
    @inproceedings{parjan2024obp,
      title = {In OBP We Trust: Verification and Validation of the M2020 On Board Planner Flight Software},
      author = {Parjan, Shreya and Gaines, Dan},
      booktitle = {IEEE Aerospace Conference},
      address = {Big Sky, Montana},
      year = {2024},
      doi = {10.48577/jpl.67lmez},
      abstract = {With a planned deployment to the Perseverance roverin October 2023, the Mars 2020 On Board Planner (OBP) flightsoftware (FSW) will allow the rover to conserve energy and au-tonomously perform more activities when surplus resources (i.e.energy, data volume, time) are available. This functionality iscritical to support efficient resource management and maximizethe productivity of the vehicle as the rover’s battery degradesover time, particularly as the Mars Sample Return mission(MSR) may rely on Perseverance to facilitate sample deliveryto the return lander. Up until OBP’s deployment, operatorshave used the Master/Sub-Master paradigm (M/SM) for allMars missions to-date, in which rigid plans are uplinked to therover based on highly conservative models of activity resourceconsumption and duration that can unnecessarily constrain on-board plans. Thus, OBP must be robust enough to merit thetrust of mission science and operations personnel and preservethe health and safety of the spacecraft. To this end, we havecompleted a thorough verification and validation (V&V) cam-paign of the first iteration of OBP that will be deployed to therover as part of new Simple Planner (SP) operations paradigm.The campaign heavily relied on software testing to ensure thatcore flight software capabilities were met and that the softwarecomprehensively covered potential edge cases. In this paper,we motivate the need for a thorough V&V campaign going intothe deployment of On Board Planner to the Perseverance rover,contextualize how OBP integrates with the rest of the vehicle’sflight software, and leverage three case studies to evidence howtesting provided critical support towards the successful infusionof onboard scheduling software on the rover.}
    }
  6. Wang, E. and Brown, J. (2007). Cassini's Test Methodology for Flight Software Verification and Operations . AIAA SPACE Conference and Exposition. Source
    BibTeX
    @inproceedings{wang2007cassini,
      title = {Cassini's Test Methodology for Flight Software Verification and Operations},
      author = {Wang, Eric and Brown, Jay},
      booktitle = {AIAA SPACE Conference and Exposition},
      year = {2007},
      url = {https://dataverse.jpl.nasa.gov/dataset.xhtml?persistentId=hdl:2014/41356}
    }
  7. Lutz, R. R. and Mikulski, I. C. (2003). Patterns of Software Defect Data on Spacecraft . NASA Office of Safety and Mission Assurance, Jet Propulsion Laboratory. Source
    BibTeX
    @techreport{lutz2003patterns,
      title = {Patterns of Software Defect Data on Spacecraft},
      author = {Lutz, Robyn R. and Mikulski, Ines Carmen},
      institution = {NASA Office of Safety and Mission Assurance, Jet Propulsion Laboratory},
      year = {2003},
      url = {https://dataverse.jpl.nasa.gov/dataset.xhtml?persistentId=hdl:2014/7132}
    }
  8. Bonitz, R., Shiraishi, L., Robinson, M., Carsten, J., Volpe, R., Trebi-Ollennu, A., Arvidson, R. E., Chu, P. C., Wilson, J. J. and Davis, K. R. (2009). The Phoenix Mars Lander Robotic Arm . IEEE Aerospace Conference. Source
    BibTeX
    @inproceedings{bonitz2009phoenix,
      title = {The Phoenix Mars Lander Robotic Arm},
      author = {Bonitz, Robert and Shiraishi, Lori and Robinson, Matthew and Carsten, Joseph and Volpe, Richard and Trebi-Ollennu, Ashitey and Arvidson, Raymond E. and Chu, P. C. and Wilson, J. J. and Davis, K. R.},
      booktitle = {IEEE Aerospace Conference},
      pages = {1-12},
      organization = {Jet Propulsion Laboratory, California Institute of Technology},
      address = {Big Sky, Montana},
      year = {2009},
      doi = {10.1109/aero.2009.4839306},
      abstract = {The Phoenix Mars Lander Robotic Arm (RA) has operated for 149 sols since the Lander touched down on the north polar region of Mars on May 25, 2008. During its mission it has dug numerous trenches in the Martian regolith, acquired samples of Martian dry and icy soil, and delivered them to the Thermal Evolved Gas Analyzer (TEGA) and the Microscopy, Electrochemistry, and Conductivity Analyzer (MECA). The RA inserted the Thermal and Electrical Conductivity Probe (TECP) into the Martian regolith and positioned it at various heights above the surface for relative humidity measurements. The RA was used to point the Robotic Arm Camera to take images of the surface, trenches, samples within the scoop, and other objects of scientific interest within its workspace. Data from the RA sensors during trenching, scraping, and trench cave-in experiments have been used to infer mechanical properties of the Martian soil. This paper describes the design and operations of the RA as a critical component of the Phoenix Mars Lander necessary to achieve the scientific goals of the mission.}
    }
  9. Lala, J. H. and Smith, T. B. I. (1983). Development and Evaluation of a Fault-Tolerant Multiprocessor (FTMP) Computer. Volume III: FTMP Test and Evaluation . Charles Stark Draper Laboratory for NASA Langley Research Center, NASA CR-166073. Source
    BibTeX
    @techreport{lala1983development,
      title = {Development and Evaluation of a Fault-Tolerant Multiprocessor (FTMP) Computer. Volume III: FTMP Test and Evaluation},
      author = {Lala, Jaynarayan H. and Smith, T. Basil, III},
      number = {NASA CR-166073},
      institution = {Charles Stark Draper Laboratory for NASA Langley Research Center},
      year = {1983},
      url = {https://ntrs.nasa.gov/citations/19850022395},
      abstract = {The experimental test and evaluation of the Fault-Tolerant Multiprocessor (FTMP) is described. Major objectives of this exercise include expanding validation envelope, building confidence in the system, revealing any weaknesses in the architectural concepts and in their execution in hardware and software, and in general, stressing the hardware and software. To this end, pin-level faults were injected into one LRU of the FTMP and the FTMP response was measured in terms of fault detection, isolation, and recovery times. A total of 21,055 stuck-at-0, stuck-at-1 and invert-signal faults were injected in the CPU, memory, bus interface circuits, Bus Guardian Units, and voters and error latches. Of these, 17,418 were detected. At least 80 percent of undetected faults are estimated to be on unused pins. The multiprocessor identified all detected faults correctly and recovered successfully in each case. Total recovery time for all faults averaged a little over one second. This can be reduced to half a second by including appropriate self-tests.}
    }