Obfuscated Gradients Give a False Sense of Security: Circumventing Defenses to Adversarial Examples

Last update: Dec 18, 2022

Overview

Obfuscated Gradients Give a False Sense of Security: Circumventing Defenses to Adversarial Examples

Above is an adversarial example: the slightly perturbed image of the cat fools an InceptionV3 classifier into classifying it as "guacamole". Such "fooling images" are easy to synthesize using gradient descent (Szegedy et al. 2013).

In our recent paper, we evaluate the robustness of nine papers accepted to ICLR 2018 as non-certified white-box-secure defenses to adversarial examples. We find that seven of the nine defenses provide a limited increase in robustness and can be broken by improved attack techniques we develop.

Below is Table 1 from our paper, where we show the robustness of each accepted defense to the adversarial examples we can construct:

Defense	Dataset	Distance	Accuracy
Buckman et al. (2018)	CIFAR	0.031 (linf)	0%*
Ma et al. (2018)	CIFAR	0.031 (linf)	5%
Guo et al. (2018)	ImageNet	0.05 (l2)	0%*
Dhillon et al. (2018)	CIFAR	0.031 (linf)	0%
Xie et al. (2018)	ImageNet	0.031 (linf)	0%*
Song et al. (2018)	CIFAR	0.031 (linf)	9%*
Samangouei et al. (2018)	MNIST	0.005 (l2)	55%**
Madry et al. (2018)	CIFAR	0.031 (linf)	47%
Na et al. (2018)	CIFAR	0.015 (linf)	15%

(Defenses denoted with * also propose combining adversarial training; we report here the defense alone. See our paper, Section 5 for full numbers. The fundemental principle behind the defense denoted with ** has 0% accuracy; in practice defense imperfections cause the theoretically optimal attack to fail, see Section 5.4.2 for details.)

The only defense we observe that significantly increases robustness to adversarial examples within the threat model proposed is "Towards Deep Learning Models Resistant to Adversarial Attacks" (Madry et al. 2018), and we were unable to defeat this defense without stepping outside the threat model. Even then, this technique has been shown to be difficult to scale to ImageNet-scale (Kurakin et al. 2016). The remainder of the papers (besides the paper by Na et al., which provides limited robustness) rely either inadvertently or intentionally on what we call obfuscated gradients. Standard attacks apply gradient descent to maximize the loss of the network on a given image to generate an adversarial example on a neural network. Such optimization methods require a useful gradient signal to succeed. When a defense obfuscates gradients, it breaks this gradient signal and causes optimization based methods to fail.

We identify three ways in which defenses cause obfuscated gradients, and construct attacks to bypass each of these cases. Our attacks are generally applicable to any defense that includes, either intentionally or or unintentionally, a non-differentiable operation or otherwise prevents gradient signal from flowing through the network. We hope future work will be able to use our approaches to perform a more thorough security evaluation.

Paper

Abstract:

We identify obfuscated gradients, a kind of gradient masking, as a phenomenon that leads to a false sense of security in defenses against adversarial examples. While defenses that cause obfuscated gradients appear to defeat iterative optimization-based attacks, we find defenses relying on this effect can be circumvented. We describe characteristic behaviors of defenses exhibiting the effect, and for each of the three types of obfuscated gradients we discover, we develop attack techniques to overcome it. In a case study, examining non-certified white-box-secure defenses at ICLR 2018, we find obfuscated gradients are a common occurrence, with 7 of 9 defenses relying on obfuscated gradients. Our new attacks successfully circumvent 6 completely, and 1 partially, in the original threat model each paper considers.

For details, read our paper.

Source code

This repository contains our instantiations of the general attack techniques described in our paper, breaking 7 of the ICLR 2018 defenses. Some of the defenses didn't release source code (at the time we did this work), so we had to reimplement them.

Citation

@inproceedings{obfuscated-gradients,
  author = {Anish Athalye and Nicholas Carlini and David Wagner},
  title = {Obfuscated Gradients Give a False Sense of Security: Circumventing Defenses to Adversarial Examples},
  booktitle = {Proceedings of the 35th International Conference on Machine Learning, {ICML} 2018},
  year = {2018},
  month = jul,
  url = {https://arxiv.org/abs/1802.00420},
}

You might also like...

Vulnerability Scanner & Auto Exploiter You can use this tool to check the security by finding the vulnerability in your website or you can use this tool to Get Shells

About create a target list or select one target, scans then exploits, done! Vulnnr is a Vulnerability Scanner & Auto Exploiter You can use this tool t

108 Dec 4, 2021

Tools to make working the Arch Linux Security Tracker easier

This is a collection of Python scripts to make working with the Arch Linux Security Tracker easier.

6 Jul 13, 2022

Writeups for wtf-CTF hosted by Manipal Information Security Team as part of Techweek2021- INCOGNITO

wtf-CTF_Writeups Table of Contents Table of Contents Crypto Misc Reverse Pwn Web Crypto wtf_Bot Author: Madjelly Join the discord server!You know how

6 Jun 7, 2021

GitHub Advance Security Compliance Action

advanced-security-compliance This Action was designed to allow users to configure their Risk threshold for security issues reported by GitHub Code Sca

121 Dec 14, 2022

evtx-hunter helps to quickly spot interesting security-related activity in Windows Event Viewer (EVTX) files.

Introduction evtx-hunter helps to quickly spot interesting security-related activity in Windows Event Viewer (EVTX) files. It can process a high numbe

116 Dec 29, 2022

Set the draft security HTTP header Permissions-Policy (previously Feature-Policy) on your Django app.

django-permissions-policy Set the draft security HTTP header Permissions-Policy (previously Feature-Policy) on your Django app. Requirements Python 3.

76 Nov 30, 2022

EMBArk - The firmware security scanning environment

Embark is being developed to provide the firmware security analyzer emba as a containerized service and to ease accessibility to emba regardless of system and operating system.

175 Dec 14, 2022

Docker is an open platform for developing, shipping, and running applications OS-level virtualization to deliver software in packages called containers However, 'security' is a top request on Docker's public roadmap This project aims at vulnerability check for such docker containers. New contributions are accepted

Docker-Vulnerability-Check Docker is an open platform for developing, shipping, and running applications OS-level virtualization to deliver software i

103 Aug 20, 2022

GitLab CI security tools runner

Common Security Pipeline Описание проекта: Данный проект является вариантом реализации DevSecOps практик, на базе: GitLab DefectDojo OpenSouce tools g

14 Dec 23, 2022

Obfuscated Gradients Give a False Sense of Security: Circumventing Defenses to Adversarial Examples

Related tags

Overview

Obfuscated Gradients Give a False Sense of Security: Circumventing Defenses to Adversarial Examples

Paper

Source code

Citation

You might also like...

Vulnerability Scanner & Auto Exploiter You can use this tool to check the security by finding the vulnerability in your website or you can use this tool to Get Shells

Tools to make working the Arch Linux Security Tracker easier

Writeups for wtf-CTF hosted by Manipal Information Security Team as part of Techweek2021- INCOGNITO

GitHub Advance Security Compliance Action

evtx-hunter helps to quickly spot interesting security-related activity in Windows Event Viewer (EVTX) files.

Set the draft security HTTP header Permissions-Policy (previously Feature-Policy) on your Django app.

EMBArk - The firmware security scanning environment

GitLab CI security tools runner

Releases(v0)

v0(Feb 4, 2018)

Owner

Anish Athalye

Bypass ReCaptcha: A Python script for dealing with recaptcha

IDA Frida Plugin for tracing something interesting.

Scans for Log4j versions effected by CVE-2021-44228

Exploit and Check Script for CVE 2022-1388

Ducky Script is the payload language of Hak5 gear.

Bandit is a tool designed to find common security issues in Python code.

S2-061 的payload，以及对应简单的PoC/Exp

CVE-2022-21907 - Windows HTTP协议栈远程代码执行漏洞 CVE-2022-21907

A Python replicated exploit for Webmin 1.580 /file/show.cgi Remote Code Execution

A Safer PoC for CVE-2022-22965 (Spring4Shell)

😭 WSOB is a python tool created to exploit the new vulnerability on WSO2 assigned as CVE-2022-29464.

spring-cloud-gateway-rce CVE-2022-22947

Cve-2021-22005-exp

Python library to prevent XSS(cross site scripting attach) by removing harmful content from data.

Python lib to automate basic QFT calculations like Wick-contractions.

DNSSEQ: PowerDNS with FALCON Signature Scheme

Automatically download all 10,000 CryptoPunk NFTs.

Getting my gitlab commit history into github

Archive-Crack - A Tools for crack file archive

Recon is a script to perform a full recon on a target with the main tools to search for vulnerabilities.