Blog · Scale

A million files is a different problem

Processing a thousand images and processing a million PDFs share a name but barely share a solution. Here’s what actually makes the big one tractable — offline, on a laptop-class machine.

Dashboard of a large-scale offline face-forensics run

Why volume changes the architecture

Hand a face-forensics tool a thousand files and it is a one-session task: point it at the folder, close the laptop, come back to results. Scale that same job to a million files and the shape of the problem changes. It is no longer a question of whether one machine can process a single item; it is a question of how the whole pipeline holds together over several hours of continuous, unattended work.

Three concerns stop mattering at small scale and become load-bearing at large scale. Can the work be spread across every core the machine has? Can it survive — and resume from — the interruptions a long run is guaranteed to hit? Can the intermediate data it produces live on disk instead of in RAM? DawaImg is designed around these three from the start, which is why the same install that comfortably handles a folder can also carry a corpus.

Parallelism and resumability

The pipeline is split so that the expensive, file-level work fans out across all available cores, while the shared state that has to stay consistent — the face groupings and the provenance links back to source pages and files — is kept serial and deterministic. Parallelizing the heavy lifting and serializing the bookkeeping is what lets a run use every core without racing on the very results it is trying to produce.

Resumability is the other half of the design. A multi-hour run crosses reboots, crashes, and other demands on the machine, and a pipeline that throws away everything when it stops is one that loses the run. Because progress is written out as files are processed, stopping a half-finished run and picking it up later costs almost nothing: you resume from the last checkpoint instead of starting from the top.

Run it at your own scale

Free, offline, and resumable — on Windows or macOS.

Disk-backed caches, not RAM hoarding

A million files generate a lot of intermediate material — crops, embeddings, and page and frame indices. The obvious approach is to hold all of it in memory, and the wrong one: it caps the size of the sets you can process to whatever fits in RAM, which is exactly the ceiling that large cases run into.

DawaImg keeps that material in disk-backed caches instead. The entries in active use stay fast, and the ones that have fallen away come back on demand. The working set is therefore bounded by what is currently being used rather than by the total size of the corpus, so memory serves the present step instead of being consumed up front by the entire set.

The numbers from our case study

Put the three ideas together and the numbers are unremarkable in the best possible way. A representative run covers about one million files on consumer-grade hardware, with the full offline pipeline — detection, grouping, and provenance — completing in roughly five to six hours.

Face detection reached about 98% recall, recovering roughly 2,000 faces across the set, and the false positives were essentially noise: two of them, both on wood textures the detector mistook for a face. None of it required a server, a cloud, or an account. It ran free and offline, with no telemetry, on a v39.0 build on a machine you could buy at any retailer.