preprint · Zenodo (CERN European Organization for Nuclear Research)
Context & Problem Face recognition models are traditionally selected based on a single metric: pair verification accuracy. However, this benchmark does not reflect real-world deployments. In production, models perform open-set identification against growing galleries under strict false-accept budgets. This methodological disconnect leads to catastrophic performance drops when models transition from benchmarks to operations. Methodology & Dataset Scope This repository contains the evaluation pipeline, benchmarks, and performance data analyzing the operational gap between verification and open-set identification. The study evaluates three prominent embedding families through a unified production scoring path: AdaFace ArcFace FaceNet Performance is mapped across: 5 standard verification benchmarks (including LFW and CPLFW). 4 open-set galleries scaling from 10 to 1,000 enrolled identities, mixed with unknown distractors. Key Findings & Replicability The dataset demonstrates that standard verification reports mislead practitioners in three distinct ways: The Amplification Effect: Small differences in headline accuracy expand drastically at strict operational thresholds. For example, FaceNet trails the leading model by 2.3% on LFW accuracy (97.50% vs 99.83%), but this deficit amplifies 27× at a False Accept Rate (FAR) of \(10^{-4}\), where True Accept Rate (TAR) plummets to 42.80% compared to 99.70%. Rank-1 Masking: Rank-1 metrics completely hide open-set failures because they ignore whether an unknown person should be rejected. At 1,000 identities, FaceNet’s rank-1 rate (84.30%) overstates its open-set accuracy (72.01%) by 12.3 points, admitting 10% of distractors. Gallery Scale Blindness: Performance gaps widen as galleries grow. The top two embedding families sit 5.0 points apart at 10 enrolled identities, but diverge to 23.6 points apart at 1,000 identities. Multi-shot enrollment mitigates but never closes this gap. The provided pipeline reproduces the deployed family's published baseline accuracy to within 0.01 points, validating the pipeline's integrity and confirming that the observed degradation is systemic rather than implementation-based.
This page summarises published work. The authoritative version sits with the publisher.
DOI: 10.5281/zenodo.21875817
Is something wrong with this record? Report it or request removal.
Discussion
Have you built on this work, tried to replicate it, or seen it applied in practice? Share what you know. Verified researchers and MARATTO™ domain experts can open a discussion, and any member can reply. Contributions are reviewed before they appear.
No discussion yet. Open the first thread.
New to MARATTO™? Create a free account.