face-lens

Two small models that run in the browser (onnxruntime-web, WebGPU or WASM) on a webcam feed:

File Input Outputs Size
student.onnx face crop, 256×256 gender, age (0-9 … 60+), emotion (7), hair color (7), eyes color (6), tongue (out or not) 17 MB
clothing.onnx crop below the chin, 224×224 style (9 garment types), pattern (solid / striped / checked / printed) 5 MB

student.json / clothing.json carry the class names, normalization and the exact crop each model was trained on. Weights are stored as fp16 and cast to fp32 inside the graph, so results are the same on WebGPU and WASM.

They are meant to be used through @receptron/face-lens (source), which also adds head direction, facial expressions (MediaPipe blendshapes), raised-finger counts and clothing colors:

import { FaceLens } from "@receptron/face-lens";

const lens = await FaceLens.create({ hands: true, clothing: true });
const r = lens.detect(video, performance.now());
// r.face?.attributes?.emotion.label, r.clothing?.style?.label, r.hands?.total, …

The library loads these files from https://huggingface.co/snakajima/face-lens/resolve/v1/.

How they were made

Knowledge distillation from Ternary Bonsai 2 27B (Apache-2.0), a vision-language model, run offline on a Mac:

  • The teacher saw each training crop once, with all questions in one prompt; its answer distribution was read directly from the logits (no generated text) and used as a soft label.
  • gender and age are trained on FairFace human labels (CC BY 4.0; ~80k faces). emotion, hair, eyes and tongue on the teacher's soft labels for ~11k FairFace faces, plus 175 openly licensed tongue-out / no-tongue photos.
  • clothing.onnx is trained on the teacher's labels for 3,890 openly licensed upper-body photos.
  • Every crop was cut exactly as the browser cuts it (MediaPipe Face Landmarker landmarks).
  • All web photos are CC BY, CC0 or public domain; ATTRIBUTION.csv lists each one with its creator, license and source page.

Evaluation

Held-out data. For teacher-labelled heads the number is agreement with the teacher, not human-verified accuracy.

Output Metric Value
gender accuracy vs FairFace labels (10,081 val faces) 95.5%
age (7 groups) accuracy vs FairFace labels 60.9%
emotion agreement with teacher 90.9%
hair color agreement with teacher 85.5%
eye color agreement with teacher 86.7%
tongue out held-out web photos 14/16 caught, 0/5 false alarms
clothing style agreement with teacher (363 val crops) 64.5%
clothing pattern agreement with teacher 85.4%

Limitations and responsible use

  • Gender, age and emotion are guesses from appearance. They are often wrong and can be biased across groups. Do not use them to make decisions about people.
  • The model deliberately does not estimate race or ethnicity.
  • FairFace's emotions are mostly neutral or happy, so sad, angry, fearful and disgusted faces are learned from few examples.
  • Eye color is the weakest attribute at webcam resolution; clothing style struggles with small, distant or layered clothing.
  • Under the EU AI Act, emotion recognition is prohibited in workplaces and education, and biometric categorisation of sensitive traits is prohibited. Check the rules where you deploy.

License

CC BY 4.0. Please credit "face-lens (snakajima)" and keep ATTRIBUTION.csv with redistributions. Training data: FairFace (CC BY 4.0) and the photos listed in ATTRIBUTION.csv. Base weights: timm MobileNetV4 (Apache-2.0). Teacher: Ternary Bonsai 2 27B (Apache-2.0).

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for snakajima/face-lens

Quantized
(1)
this model

Dataset used to train snakajima/face-lens