research-article

Context-Aware Fuzzing for Robustness Enhancement of Deep Learning Models

Authors:

Haipeng Wang,

Zhengyuan Wei,

Qilin Zhou,

Wing-Kwong ChanAuthors Info & Claims

ACM Transactions on Software Engineering and Methodology, Volume 34, Issue 1

Article No.: 8, Pages 1 - 68

https://doi.org/10.1145/3680464

Published: 31 December 2024 Publication History

Get Access

Abstract

In the testing-retraining pipeline for enhancing the robustness property of deep learning (DL) models, many state-of-the-art robustness-oriented fuzzing techniques are metric-oriented. The pipeline generates adversarial examples as test cases via such a DL testing technique and retrains the DL model under test with test suites that contain these test cases. On the one hand, the strategies of these fuzzing techniques tightly integrate the key characteristics of their testing metrics. On the other hand, they are often unaware of whether their generated test cases are different from the samples surrounding these test cases and whether there are relevant test cases of other seeds when generating the current one. We propose a novel testing metric called Contextual Confidence (CC). CC measures a test case through the surrounding samples of a test case in terms of their mean probability predicted to the prediction label of the test case. Based on this metric, we further propose a novel fuzzing technique Clover as a DL testing technique for the pipeline. In each fuzzing round, Clover first finds a set of seeds whose labels are the same as the label of the seed under fuzzing. At the same time, it locates the corresponding test case that achieves the highest CC values among the existing test cases of each seed in this set of seeds and shares the same prediction label as the existing test case of the seed under fuzzing that achieves the highest CC value. Clover computes the piece of difference between each such pair of a seed and a test case. It incrementally applies these pieces of differences to perturb the current test case of the seed under fuzzing that achieves the highest CC value and to perturb the resulting samples along the gradient to generate new test cases for the seed under fuzzing. Clover finally selects test cases among the generated test cases of all seeds as much as possible and with a preference to select test cases with higher CC values for improving model robustness. The experiments show that Clover outperforms the state-of-the-art coverage-based technique Adapt and loss-based fuzzing technique RobOT by 67%–129% and 48%–100% in terms of robustness improvement ratio, respectively, delivered through the same testing-retraining pipeline. For test case generation, in terms of numbers of unique adversarial labels and unique categories for the constructed test suites, Clover outperforms Adapt by \(2.0\times\) and \(3.5\times\) and RobOT by \(1.6\times\) and \(1.7\times\) on fuzzing clean models, and also outperforms Adapt by \(3.4\times\) and \(4.5\times\) and RobOT by \(9.8\times\) and \(11.0\times\) on fuzzing adversarially trained models, respectively.

References

[1]

GitHub, Inc. 2017. Fashion-mnist. Retrieved from https://github.com/zalandoresearch/fashion-mnist

Abstract

References

Index Terms

Recommendations

Neuron Semantic-Guided Test Generation for Deep Neural Networks Fuzzing

Fuzzing for Deep Learning Models

Cerebro: context-aware adaptive fuzzing for effective vulnerability detection

Comments

Information

Published In

Publisher

Publication History

Check for updates

Author Tags

Qualifiers

Funding Sources

Contributors

Other Metrics

Bibliometrics

Article Metrics

Other Metrics

Citations

Login options

Full Access

View options

PDF

eReader

Full Text

Share

Share this Publication link

Share on social media

Affiliations