Scinovex
article Open AccessTop 1% cited

BadNets: Evaluating Backdooring Attacks on Deep Neural Networks

IEEE Access · 2019 · Vol. 7 · pp. 47230–47244
Tianyu GuKang LiuBrendan Dolan-GavittSiddharth Garg

Abstract

Deep learning-based techniques have achieved state-of-the-art performance on a wide variety of recognition and classification tasks. However, these networks are typically computationally expensive to train, requiring weeks of computation on many GPUs; as a result, many users outsource the training procedure to the cloud or rely on pre-trained models that are then fine-tuned for a specific task. In this paper, we show that the outsourced training introduces new security risks: an adversary can create a maliciously trained network (a backdoored neural network, or a BadNet) that has the state-of-the-art performance on the user's training and validation samples but behaves badly on specific attacker-chosen inputs. We first explore the properties of BadNets in a toy example, by creating a backdoored handwritten digit classifier. Next, we demonstrate backdoors in a more realistic scenario by creating a U.S. street sign classifier that identifies stop signs as speed limits when a special sticker is added to the stop sign; we then show in addition that the backdoor in our U.S. street sign detector can persist even if the network is later retrained for another task and cause a drop in an accuracy of 25% on average when the backdoor trigger is present. These results demonstrate that backdoors in neural networks are both powerful and-because the behavior of neural networks is difficult to explicate-stealthy. This paper provides motivation for further research into techniques for verifying and inspecting neural networks, just as we have developed tools for verifying and debugging software.

Adversarial Robustness in Machine LearningAnomaly Detection Techniques and ApplicationsAdvanced Malware Detection TechniquesBackdoorComputer scienceTraffic sign recognitionAdversaryArtificial neural networkArtificial intelligenceClassifier (UML)Deep neural networksDeep learningMachine learning

Funding

  • National Science Foundation
Citations
1,092
FWCI
49.49
field-weighted impact
References
62
Percentile
100%
vs. same field & year
Citations per year
References
Deep learning in neural networks: An overview
Neural Networks · 2014 · 17,774 citations
ImageNet classification with deep convolutional neural networks
Communications of the ACM · 2017 · 75,550 citations
Citation Network

How this paper connects to the literature. Drag to explore, click any node to open that paper.

BadNets: Evaluating Backdooring Attacks on Deep Neural Networks · Scinovex