A preprint steers deep neural network navigation testing toward regions where an explainability method called Integrated Gradients flags the model as least confident, finding more transferable failures than blind search on a LeoRover small wheeled
Researchers have developed an explainability-guided approach to test the robustness of deep neural network (DNN) controllers in robotic systems, achieving substantially higher failure-detection rates than conventional search methods.
The technique, described in a preprint, uses Integrated Gradients—an explainability method that attributes a model's prediction to its input features—to guide an evolutionary search toward image regions where the DNN is least confident. On a simulated LeoRover robot, the approach reached a 70.0% median success rate in finding transferable failures, compared to 53.85% for unguided search, according to the paper (arxiv.org/abs/2610.06862).
The researchers first clustered representative images by visual appearance and behavioral output, then applied sparse perturbations guided by aggregated Integrated Gradients. A multi-image aggregation step alone improved success from 50.0% to 57.5%; adding XAI guidance pushed it to 70.0%. The Pareto hypervolume—a measure of solution quality across multiple objectives—lifted from 0.65 to 0.73.
To validate real-world relevance, the team transferred the failures discovered in simulation to a physical LeoRover. The transfer achieved 0.95 precision and 0.67 recall for failure detection, with a failure-time correlation of 0.617. The authors note that while simulation effectively identifies transferable failures, it is not a substitute for physical validation.
The work highlights the potential of explainability-guided robustness testing for DNN-controlled cyber-physical systems, though the results are limited to a single platform and robot model.