Publication
Catching a Confident Lie
Josh Harguess, J. Takeshi Miclat, and Julian Raheema. "Using image quality metrics to identify adversarial imagery for deep learning networks." Proc. SPIE 10199, Geospatial Informatics, Fusion, and Motion Video Analytics VII, 1019907. 2017.
Imagine that somewhere there's a security guard falling asleep in front of a collection of monitors, trusting an AI system that reports "all clear, 99% confidence."
That confidence number is the whole basis for trusting the system, and it turns out you can build an image specifically to make an AI report exactly that, with total certainty, while showing it something that looks like a static screen to a human.
That's not a hypothetical. In 2017, working with the Navy's research arm, we built and tested images designed to do exactly this, and confirmed a deep learning classifier really would identify a smear of noise as "a strawberry" with near-total confidence.
The scary part isn't that the model was wrong. Models are wrong sometimes. The scary part is it was sure.
So if a client asks for confidence scores they can rely on, this is the fine print worth knowing: confidence is a property of the model, and it's exactly the thing an attack is built to exploit. If you can't trust the model to know when it's being fooled, you need something that doesn't depend on the model's own judgment at all.
Working with Josh Harguess and Julian Raheema, published through SPIE, we tested whether a no-reference image quality metric, Blind/Referenceless Image Spatial Quality Evaluator (BRISQUE), could flag adversarial imagery before it ever reached the classifier.
The BRISQUE metric works by acting in the same sort of way a blacklight does:
- The "blacklight" is shone on parts of the darkened room, by making it greyscale, at a time via the localized Gaussian window. This just means what part of the image you're looking at.
- Identify what is normal and not "shining" in the blacklight's light. This is done via the local mean, average brightness of the window, and local variance, how much the contrasts vary from that average.
You then level all the textures by:
- Stripping the lighting, meaning that super bright and super dark parts of the photo are canceled out so the photo can be judged on equal grounds.
- Standardizing the contrast, meaning that chaotic 80's-style furniture colors are dampened and minimalist corporate furniture is boosted.
In the end, you can see a map of raw texture that has been leveled to the same playing field. There's now a clear way for our blacklight to tell us when something is weird.
If we plot every pixel on a histogram for a large corpus of images, ImageNet and MNIST, we can see that the curve always looks a certain way: perfectly symmetrical and shaped similarly to each other.
We can then create adversarial images by either:
- Programmatically changing pixels against random noise until a deep neural network classifier is fooled.
- Using a generative network to apply complex patterns based on how a neural network "learns" to classify the image.

If we do the same histogram plot for these resulting adversarial images, the chart would look different:
- Compression artifacts, JPEG-type files that take up less space than their PNG counterpart, distort the symmetry.
- Blur squashes the curve from the sides to make it look taller and narrower.
- White noise flattens the curve down and makes it look wider.


So we were able to programmatically detect adversarial images, static, blur, camera occlusions, and image or video file corruption. This means we can:
- Protect ML systems from being fooled by adding a layer of validation.
- Know how to direct human manual inspection and review of footage that could be hundreds and thousands of hours long by identifying these anomalies.
- Quickly alert someone when a security camera has something covering it.
This system formed one of the earliest parts of my AI product DNA: A model's confidence can't check itself. The check has to come from somewhere that doesn't share its blind spots.
The choice that made the approach work: don't try to make the classifier more robust. Check something upstream of it that has no stake in the classifier's opinion.
Every layer of that system, the classifier itself, its confidence score, was exactly what the attack was designed to compromise. Anything derived from the model's own judgment was untrustworthy by construction. BRISQUE worked as a check precisely because it was blind to classification entirely, it only asked "does this look like a normal photograph," a question adversarial optimization wasn't targeting.
It worked well on ImageNet, cleanly separating real from adversarial images. It worked less well on MNIST, and the paper is direct about why: BRISQUE is trained on natural photographs, and handwritten digits aren't natural photographs. The metric's own assumptions didn't transfer to a domain it wasn't built for.
A model's own confidence can't be the thing that tells you whether to trust it, because confidence is a property of the model, and the model is exactly what's in question. You need a check that doesn't share the model's blind spots, and you need to know where that check's own assumptions stop holding.
You can see it in nearly every AI case study on this site. Everything is built on the idea that validation independent of the model producing the output, and honesty about where any given safeguard's assumptions run out, is what makes an AI system trustworthy.