arXiv:2608.21486nvidia/nemotron-3.5-lightning-30b-a3bAugust 21, 2026

EXPL-FR: Explaining Face Recognition Models via Vision-Language AlignmentExplained for Beginners

Guray Ozgur, Mustafa Efe Tamyapar, Naser Damer +1 more

Computer Vision and Pattern Recognition

Abstract

Deep face recognition (FR) models reach near-saturated accuracy but remain opaque: a practitioner cannot ask which semantic attributes a similarity score relied upon. EXPL-FR answers this inside the FR model's own embedding space. A lightweight adapter aligns a vision-language model's (VLM) image encoder with the frozen FR space, trained on face images alone and never on text. Because the VLM's encoders share one space, the same adapter applies to the text encoder, turning 978 attribute prompts in 22 categories, also extendable, into FR-space anchors at no extra cost. We do not assume this transfer works: a face-verification protocol measures it, and an ablation changing only the adapter isolates its contribution. Not every concept survives, because an FR model earns its invariances by discarding the factors it must verify identities across. A label-free detectability measure compares each concept's separability in FR space against the VLM space, and the 100 most detectable form the model's readable semantic signature, which separates identities better than the full vocabulary. We cover four FR backbones and two VLM encoders, EXPL-FR needs no architecture access, and supports identity-level, per-image, and differential explanations. We benchmark attribute-level auditing under three supervision settings, human labels (current practice), VLM pseudo-labels, and our fully prompt-driven audit, against real verification behavior. With no labels, the prompt-driven audit ranks four FR models by their measured per-ethnicity RFW errors and ranks controlled attribute changes by their true verification cost.

Explaining Face Recognition Models via Vision-Language Alignment

1. The Problem: The Silence of the Score

Imagine you are a border agent or a phone unlock system designer. A face recognition (FR) model compares two images and spits out a single number: a similarity score. If the score passes a threshold, the system says “yes, these are the same person”; if not, “no.” Simple. But here is the frustration: that number is a black box. You get a yes or no, but you have absolutely no idea why the model made that decision.

Maybe the model matched on the nose. Maybe it was confused by the lighting. Maybe it focused on a facial tattoo that has nothing to do with identity. In high-stakes environments—law enforcement, banking, device security—this opacity is unacceptable. Regulators and users alike are demanding transparency. We need to know: Which semantic attributes drove this similarity score?

Prior attempts to explain FR models have been clumsy. Some try to visualize which pixels the model looks at (saliency maps), but that tells you where it looked, not what it recognized. Other methods require training a "concept detector" on top of the model, which is impossible if you are using a deployed, closed-source system. You can't just tack on a new layer of analysis to a model you don't have access to.

This paper, EXPL-FR, cuts through this Gordian knot. It asks: Can we map the model's internal reasoning onto a vocabulary that humans actually understand, without ever changing the model itself?

2. How It Works: The "Translation Adapter" Analogy

The core innovation of EXPL-FR is a deceptively simple mechanism: a lightweight "adapter."

The Setup

Think of the Face Recognition model as a resident who speaks only a obscure, internal language of numbers (embeddings). The Vision-Language Model (VLM)—like CLIP—is a bilingual citizen who knows both images and text. The gap between them is the "modality gap": the VLM’s image encoder and text encoder live in a shared space, but the FR model lives in its own silo.

The authors freeze both the FR model and the VLM. They then introduce a tiny translator: a 4-layer Multi-Layer Perceptron (MLP) with only about 1.18 million parameters. That is tiny compared to the billions of parameters in the models it connects.

The Training Process (The "Image-to-Image" Phase)

The authors show the adapter millions of face images. For each image, they get two embeddings: one from the VLM’s image encoder and one from the FR encoder. The adapter learns to twist the VLM embedding so that it points in the same direction as the FR embedding. Crucially, the adapter never sees any text. It is trained purely on image-to-image alignment.

The Magic Trick (Zero-Shot Transfer to Text)

This is the clever part. Once the adapter is trained to map VLM images to FR space, the authors simply point the adapter at text.

Because the VLM’s own image and text encoders already share a space (they were trained together during the VLM's original training), the same adapter that worked for images suddenly works for text prompts. A prompt like "A photo of a person wearing glasses" gets embedded by the VLM, run through the adapter, and—presto—it lands as an anchor in the FR model's space.

The authors call these anchors the "semantic signature." When they take a face image and compute its similarity to these 978 text-derived anchors, they get a vector of numbers. These numbers tell you exactly what the FR model "thinks" about that face in terms of hair color, ethnicity, age, and so on.

The Filter: Which Concepts Survive?

Not every prompt survives the trip. The authors argue that FR models earn their invariances by discarding things that don't matter for identity—like lighting or background. So, if you translate a "sunny day" prompt through the adapter, the FR model might just ignore it because it verified identity under a million different lighting conditions.

To handle this, the authors developed a "label-free detectability measure." They compare how separable a concept is in the VLM's space versus the FR model's space. They keep the top 100 most detectable concepts. These form the model's "readable semantic signature."

Analogy: Imagine you have a dog that can recognize its owner even when wearing a hat or standing in the rain (invariances). If you ask the dog to identify "wearing a hat," it might shrug because the hat doesn't change who the owner is. EXPL-FR figures out which concepts the dog can still identify, and those become its "semantic signature."

3. Key Results & Benchmarks: What the Numbers Say

The results are compelling. The authors benchmarked four different Face Recognition backbones and two different Vision-Language Models. The "translation adapter" consistently worked, though with varying fidelity.

Faithfulness of the Bridge

In Table 1 of the paper, they measure how well the adapter preserves the FR model's identity verification capability.

  • The Upper Bound: If you verify identity directly in the FR space, you get near-perfect accuracy (around 97-99%).
  • The Bridge: Using the adapter to map VLM features into FR space drops accuracy, but not catastrophically. For the best setup (AdaFace/ViT-B with CLIP), the verification accuracy dropped to roughly 95.7%. This suggests the adapter is a faithful, though lossy, proxy.

The "Detectability" Advantage

The most striking result concerns the "100 most detectable attributes." The authors tested whether using only these 100 concepts (the semantic signature) is better than using all 978 prompts or random sets of 100.

The result? The 100 most detectable attributes separate identities better than the full vocabulary of 978.

This is counter-intuitive. You would think more information is better. But the paper explains this by the "discarding" mechanism: the full vocabulary includes many concepts the FR model has actively discarded (like lighting). Including those "noisy" concepts dilutes the signal. By filtering to the top 100, you strip away the noise and highlight what the model actually cares about for identity.

Audit Results: Supervision Levels

The authors also audited the models to see if their explanations matched real-world verification behavior. They compared three levels of supervision:

  1. Human Labels (Current Practice): Using annotated attributes (like those in CelebA) to build axes.
  2. VLM-Pseudo-Labels: Using the VLM's own guesses about attributes.
  3. Prompt-Driven (Ours): The fully zero-shot method described above.

The results, validated against difficult benchmarks like RFW (racial faces in the wild) and GAN-Control, showed that the prompt-driven audit (which requires zero labeled data) performed remarkably well. It could rank FR models by their per-ethnicity error rates and quantify the true verification cost of attribute changes.

Translating a benchmark: On the RFW benchmark (which tests racial bias), the prompt-driven audit could rank four different FR models by their measured error rates. This means a practitioner could use this method to diagnose which model is most biased against which ethnicity, simply by showing the model some text prompts, without needing any ground-truth ethnicity labels.

4. Why It Matters: Key Takeaways

This paper matters because it demystifies the "black box" of face recognition with zero operational cost.

  • Transparency without Retraining: You don't need to fine-tune the FR model or collect new data. You just need the model's embedding (which is usually already computed for the verification task) and the adapter. This makes the method immediately applicable to existing deployed systems.
  • Attribute-Level Auditing: The authors show you can audit a model for specific traits. Want to know if a model is relying too heavily on skin color? The semantic signature will highlight "ethnicity" as a strong signal. Want to check if makeup is confusing the model? Look at the "makeup" coordinate. This enables a new kind of model debugging.
  • Zero-Shot Extensibility: Because the vocabulary is just text prompts, you can instantly add new attributes. Want to check for "wearing a surgical mask"? Write a prompt. The adapter handles the rest. No re-labeling of datasets is required.
  • The "Discarding" Insight: The most profound takeaway is the realization that FR models intentionally discard information. The paper provides a quantitative measure of this. When you read the "semantic signature," you aren't just seeing what the model can see; you are seeing what the model chose to preserve to verify identity. This reframes explainability: an explanation isn't just a list of features; it's a map of the model's invariances.

Summary

EXPL-FR offers a practical, elegant solution to the opacity of face recognition. By introducing a tiny adapter trained on images, it leverages the existing multilingual capabilities of Vision-Language Models to translate FR embeddings into human-readable text. It filters out the noise, keeping only the "identity-discriminating" attributes. The result is a semantic signature that explains model decisions, enables auditing for bias, and requires zero retraining or labeled data. For anyone deploying FR systems, this paper provides the tools to finally open the black box.

Want to understand AI papers like this from scratch?

Follow the free AI Learning Roadmap →