The setting most reviewers get wrong
Similarity isn’t a dial; it’s a tradeoff between catching every relevant face and drowning in near-misses. Picking it well — and checking it — is a review skill worth practicing.
High recall vs. high precision
Every similarity threshold sits in the middle of a tradeoff. Recall is how much of what matters actually shows up: at a loose setting, few true matches fall outside the results. Precision is how trustworthy the list is: at a tight setting, nearly everything returned is a genuine match rather than a lookalike. A face that is just barely similar is either in the list or out of it, and the threshold decides which.
A loose threshold buys confidence that you didn’t miss anything, and costs time: the list grows with different people, an unlucky angle, harsh lighting, partial faces — near-misses you filter by hand. A tight one keeps the list short and defensible, but the face you were looking for may land just below the line and go unseen. Neither pole is “wrong”; a case that loses nothing by seeing extra faces wants a very different level from one whose artifact has to stand up to scrutiny.
Dial it in on real data
Free, offline similarity search. Windows & macOS.
Choosing a starting threshold
Don’t start at a guess — start at the middle and let the results tell you which way to move. Run the search at the default level and look at the list: are the top entries clearly the people you expect, or is it padded with plausible-looking strangers? Short and clearly right, and you may be too tight; a long roll of near-misses, and you’re probably too loose. Adjust in small steps, re-running, and note how the list changes.
Two things make this easier. First, similarity in DawaImg is probabilistic, not binary: each face carries a measured degree of similarity, and the threshold is a filter over those degrees — so a nudge changes the list, but you can still see how close each excluded face was. Second, the setting is per-review, not global: every search can use the level that fits it. And because DawaImg is free, offline, and needs no account, re-running a search to test a level costs nothing.
Sanity-checking: the small-group test
Before you trust a threshold on a full review, verify it against something you already know. Pick a small set with a known answer: a person you’re certain of, a pair you know are different people, or someone you know should not match. Run the search at the current level. The known match should come in; the known non-match should stay out. It takes under a minute, and it catches most mis-set thresholds before they quietly shape a long review.
A second check is stability: run the same search twice and confirm the list looks the same. If borderline entries keep shuffling around your threshold, the group sits right at the edge of that level — which is information, not a malfunction. Loosen to cast a wider net, or tighten and accept the risk, deliberately.
Match the threshold to the review goal
The “best” level is the one that fits the question. Looking for every appearance of one person across a large corpus — a device export, a document release? A looser threshold is the safer default: an extra frame to dismiss beats a missed one you never see. Building a tight, presentable list of a small group? A tighter level is worth the risk, because every entry has to carry weight.
So the practice is: pick a level for the goal, sanity-check it on a small known group, then let the list decide the next small adjustment. Repeat until the results are as defensible as the review needs them to be. That loop — choose, check, adjust — is the whole skill, and you can put it to work on your own review right now, offline, in DawaImg 39.0.