Blog · Research

The value of watching it fail

The best way to understand a face system is to make it surprise you. Detection labs and defensive masking exist to document how and why these tools miss — so the people relying on them know where the edges are.

Code on a laptop screen, representing defensive evaluation of face systems

Robustness is a question, not a badge

“Detection under hard conditions” is usually sold as a capability: low light, partial occlusion, small faces. But a capability claim is only as good as the evaluation behind it — and most of the time, that evaluation is invisible. Robustness becomes a question the moment you start asking what gets missed, how often, and where. A detector that finds faces in 99 percent of a curated folder can still hand you surprises when the folder is a week of security-camera exports, a drive full of scanned scans, or a mix of both.

That distinction is why we treat robustness as something to measure, not a label to display. DawaImg is free and offline, which makes this unusually easy: you can build your own test conditions on your own machine and watch the same version — v39.0 right now — behave differently across them.

Run the labs yourself

Free, offline evaluation on Windows & macOS.

Detectability vs. recognizability

Two properties are often muddled together, and they fail in different ways. Detectability is about whether a face is present at all: enough pixels, enough contrast, enough of an unoccluded face left in the frame. Recognizability — here, the ability of the system to group and compare faces consistently — is a separate bar. Even a fully visible face can be undetectable in a dark corridor frame, while a blurry, partial face can still register.

In one case study we ran, a set of about 2,000 faces produced only 2 false positives — both on wood-grain texture. That is a small miss rate for a detector, and exactly the kind of result you want to see on your own data, because “two false positives” is not the same conclusion for a folder of portraits as for a folder of surveillance stills.

Masks that reduce detection — and what that teaches

Defensive masking is not about hiding a face from a human reviewer. It is a calibration instrument. If you apply a mask and the detector stops finding that face reliably, you have learned two things at once: the detector was sensitive to the region you masked, and a face is less robust than a detector's marketing suggests. If the mask does nothing, you have learned that the detector's confidence was never based on the pixels you thought mattered.

Either outcome is a data point. That is the whole discipline: run the same image, change one variable, write down what changed — and do it on a machine where nothing about the experiment has to leave your control.

Doing this responsibly

Failure-mode work touches real people's faces, so the responsible defaults are unglamorous and non-negotiable: work only on media you already lawfully hold, keep results on your own machine, and never treat a grouping as an identity. DawaImg ships with none of the countermeasures built in — no accounts to link results to, no telemetry to expose what you tested, no cloud copies of your case material.

That is not a feature list; it's a lab environment. Run the detector, run the mask, look at what got missed. The people who build and rely on these systems are in a better position when they have watched it fail, not when they are hoping it is perfect.