Grouping, not knowing
The most common way face tools mislead is by implying they know who someone is. DawaImg is built on the opposite premise: it organizes appearances, and leaves knowing to the person doing the review.
Clustered, not identified
Face clustering is a geometric act, not a recognition act. DawaImg computes a compact numeric fingerprint of every face it detects, then draws a line between fingerprints that are close and calls everything on the same side a cluster. That is the whole of it: “these appearances are similar by measure.” It is not “these are the same human being,” and it is far from “this is [name].” One person can appear in several clusters at once, and a single face can split across two.
That limitation is the premise DawaImg is built on. The software does the grouping; the person doing the review does the knowing. Nothing in the interface is trying to read a name back to you, and nothing should — because it does not actually have one to read.
What a “person” is — and isn’t
When we say DawaImg works with “people,” we mean a cluster of appearances with a label you attach. A label is a note you add as the case gives you more context — “driver, second seat,” “unknown, blue shirt,” “the person in both of these.” The cluster answers “which of these images belong together.” It never answers “who is that.” Those are different questions, and confusing them is how face tools start to mislead.
So what a “person” is not: a verified identity. It is not a match against a database you did not supply, and it is not a verdict. Clustering is probabilistic at its core, and its boundaries are a similarity threshold that is deliberately configurable, because the right value depends on your footage, your angles, and your tolerance for error. DawaImg hands you the grouping and leaves the meaning to you.
Organize your next dataset
Free, offline clustering on Windows & macOS.
Re-cluster and relabel as the case unfolds
Real cases are iterative, and clustering should be too. Run it early, when you still know little, and it hands you a map you can actually look at: six clusters spread across four source files, instead of a wall of frames. Then the map shrinks as you read — you can relabel a cluster “this is the driver,” merge two that turn out to be one person, or split one that quietly mixed two similar-looking faces.
None of it has to be permanent. Adjust the threshold, re-run, and the clusters re-form around your new setting without ever touching the underlying media or the provenance records that tie each face to its source. Learning something new is not a reason to start over; it is a reason to re-cluster. A tool that keeps your changing model is a tool that can keep pace with a real case.
Where grouping genuinely helps
Held to what it actually is, clustering is exactly the work it is good at: making “who appears, and where” navigable. It turns a flat mass of frames into a small set of groups, each of which you can open, count, and trace back to the file, frame, or page it came from. It does not hand you an answer about identity; it hands you a structure good enough to work from, and that structure is defensible.
It runs the way you would want it to on sensitive data — free, offline, on your own Windows or macOS machine, with no accounts and no telemetry. In v39.0 today it will not try to turn your clusters into a registry, because it has no registry to turn them into. It will cluster, label along, and leave knowing to the person at the screen. That restraint is the entire point.