license: mit
license_link: LICENSE
library_name: openvino
pipeline_tag: image-classification
tags:
- openvino
- intel
- face-detection
- face-reidentification
- edge-ai
- metro
- dlstreamer
language:
- en
Facial Recognition
| Property | Value |
|---|---|
| Category | Face Detection + Re-Identification |
| Base Model | face-detection-adas-0001 + face-reidentification-retail-0095 (Open Model Zoo) |
| Source Framework | Caffe / PyTorch (Open Model Zoo) |
| Supported Precisions | FP32, FP16 |
| Inference Engine | OpenVINO |
| Hardware | CPU, GPU, NPU |
| Detected Class(es) | Human faces (detection) + 256-d face embeddings (re-identification) |
Overview
Facial Recognition is a Metro Analytics use case that detects human faces in images and video and computes a 256-dimensional embedding vector for each face, enabling enrollment, search, and identification against a known gallery.
The pipeline composes two Intel Open Model Zoo models:
- face-detection-adas-0001 -- an SSD-based face detector optimized for automotive and surveillance cameras (FP16, 384x672 input).
- face-reidentification-retail-0095 -- a compact CNN that maps a cropped face to a 256-d embedding; cosine similarity between embeddings determines identity.
These models are well-tested with OpenVINO Runtime and Intel DLStreamer's
gvadetect + gvaclassify pipeline.
Typical Metro deployments include:
- Access Control -- match employees or authorized personnel against an enrollment gallery.
- VIP Identification -- recognize known individuals in a crowd.
- Search and Forensics -- find a person of interest across multiple camera feeds.
- Attendance Tracking -- log when enrolled individuals enter or leave a facility.
Privacy Note: Facial recognition involves biometric data. Ensure your deployment complies with applicable privacy regulations (GDPR, BIPA, etc.) and has proper consent mechanisms in place.
Prerequisites
- Python 3.11+
- Install OpenVINO (latest version)
- Install Intel DLStreamer (latest version)
Create and activate a Python virtual environment before running the scripts:
python3 -m venv .venv --system-site-packages
source .venv/bin/activate
Note: The
--system-site-packagesflag is required so the virtual environment can access the system-installed OpenVINO and DLStreamer Python packages.
Getting Started
Download Models
Run the provided script to download the face detection and re-identification models from the Open Model Zoo:
chmod +x export_and_quantize.sh
./export_and_quantize.sh
The script performs the following steps:
- Installs
openvino. - Downloads
face-detection-adas-0001(FP16) into./intel/face-detection-adas-0001/FP16/. - Downloads
face-reidentification-retail-0095(FP16) into./intel/face-reidentification-retail-0095/FP16/. - Downloads a sample test video (
test_video.mp4).
OpenVINO Sample
The sample below runs recognition on the sample video. It detects every face,
computes a 256-d embedding, and matches it against a gallery of previously seen
people. Each new person is enrolled and assigned a numeric ID; when the same
person is seen again, the gallery returns their existing ID. Every face is
annotated with its ID <n>, and the result is saved to output_openvino.mp4.
Change the device string to run on CPU, GPU, or NPU.
import cv2
import numpy as np
import openvino as ov
DETECTION_MODEL = "intel/face-detection-adas-0001/FP16/face-detection-adas-0001.xml"
REID_MODEL = "intel/face-reidentification-retail-0095/FP16/face-reidentification-retail-0095.xml"
INPUT_VIDEO = "test_video.mp4"
CONF_THRESHOLD = 0.6
MATCH_THRESHOLD = 0.5
core = ov.Core()
# Change device to "GPU" or "NPU" to run on integrated GPU or NPU.
det_model = core.compile_model(core.read_model(DETECTION_MODEL), "CPU")
reid_model = core.compile_model(core.read_model(REID_MODEL), "CPU")
det_input = det_model.input(0)
det_h, det_w = det_input.shape[2], det_input.shape[3]
reid_input = reid_model.input(0)
reid_h, reid_w = reid_input.shape[2], reid_input.shape[3]
def detect_faces(img):
h0, w0 = img.shape[:2]
blob = cv2.resize(img, (det_w, det_h))
blob = blob.transpose(2, 0, 1)[np.newaxis, ...].astype(np.float32)
detections = det_model([blob])[det_model.output(0)][0][0]
boxes = []
for det in detections:
if float(det[2]) < CONF_THRESHOLD:
continue
x1 = max(0, int(det[3] * w0))
y1 = max(0, int(det[4] * h0))
x2 = min(w0, int(det[5] * w0))
y2 = min(h0, int(det[6] * h0))
if x2 > x1 and y2 > y1:
boxes.append((x1, y1, x2, y2))
return boxes
def get_embedding(img, bbox):
x1, y1, x2, y2 = bbox
crop = img[y1:y2, x1:x2]
blob = cv2.resize(crop, (reid_w, reid_h))
blob = blob.transpose(2, 0, 1)[np.newaxis, ...].astype(np.float32)
emb = reid_model([blob])[reid_model.output(0)].flatten()
return emb / np.linalg.norm(emb)
# Gallery of (numeric_id, embedding). recognize() returns an existing ID for a
# known face or enrolls a new one, keeping each person's ID stable over time.
gallery = []
next_id = 1
def recognize(embedding):
global next_id
best_index, best_sim = -1, 0.0
for index, (_, gallery_emb) in enumerate(gallery):
sim = float(np.dot(embedding, gallery_emb))
if sim > best_sim:
best_sim, best_index = sim, index
if best_sim >= MATCH_THRESHOLD:
person_id, gallery_emb = gallery[best_index]
# Blend the embedding into the gallery entry to stay robust to pose.
updated = 0.9 * gallery_emb + 0.1 * embedding
gallery[best_index] = (person_id, updated / np.linalg.norm(updated))
return person_id
person_id = next_id
next_id += 1
gallery.append((person_id, embedding))
print(f"Enrolled ID {person_id}")
return person_id
cap = cv2.VideoCapture(INPUT_VIDEO)
fps = cap.get(cv2.CAP_PROP_FPS) or 12
frame_w = int(cap.get(cv2.CAP_PROP_FRAME_WIDTH))
frame_h = int(cap.get(cv2.CAP_PROP_FRAME_HEIGHT))
writer = cv2.VideoWriter(
"output_openvino.mp4", cv2.VideoWriter_fourcc(*"mp4v"), fps, (frame_w, frame_h))
while True:
ok, frame = cap.read()
if not ok:
break
for bbox in detect_faces(frame):
person_id = recognize(get_embedding(frame, bbox))
x1, y1, x2, y2 = bbox
cv2.rectangle(frame, (x1, y1), (x2, y2), (0, 255, 0), 2)
cv2.putText(frame, f"ID {person_id}", (x1, max(15, y1 - 8)),
cv2.FONT_HERSHEY_SIMPLEX, 0.6, (0, 255, 0), 2)
writer.write(frame)
cap.release()
writer.release()
print(f"Total identities recognized: {len(gallery)}")
print("Saved: output_openvino.mp4")
Device targets:
"CPU"-- default, works on all Intel platforms."GPU"-- Intel integrated or discrete GPU."NPU"-- Intel NPU; face-detection-adas-0001 FP16 is NPU-compatible.
Expected Output
DLStreamer Sample
The pipeline below runs the face detector via gvadetect and the
re-identification model via gvaclassify on the video. Frames are pulled through
an appsink, where each face's embedding is matched against a gallery to assign
a stable numeric ID (new people are enrolled, returning people keep their ID).
Every face is annotated with its ID <n> and the result is saved to
output_dlstreamer.mp4.
Notes on running this sample:
Export
PYTHONPATHso the DLStreamer Python modules (gi,gstgva) are importable:source /opt/intel/openvino_2026/setupvars.sh source /opt/intel/dlstreamer/scripts/setup_dls_env.sh export PYTHONPATH=/opt/intel/dlstreamer/python:\ /opt/intel/dlstreamer/gstreamer/lib/python3/dist-packages:${PYTHONPATH:-}The re-identification embedding is attached as a tensor on each face's region-of-interest metadata. Convert the stream to
BGRbeforegvadetect/gvaclassifyso a downstream format conversion does not strip those tensors before theappsinkreads them.
import gi
gi.require_version("Gst", "1.0")
from gi.repository import Gst
Gst.init([])
import numpy as np
import cv2
from gstgva import VideoFrame
INPUT_VIDEO = "test_video.mp4"
OUTPUT_VIDEO = "output_dlstreamer.mp4"
DETECTION_MODEL = "intel/face-detection-adas-0001/FP16/face-detection-adas-0001.xml"
REID_MODEL = "intel/face-reidentification-retail-0095/FP16/face-reidentification-retail-0095.xml"
# For CPU: change "GPU" to "CPU". For NPU: change "GPU" to "NPU".
DEVICE = "GPU"
DET_THRESHOLD = 0.6
MATCH_THRESHOLD = 0.5
# Gallery of (numeric_id, embedding). recognize() returns an existing ID for a
# known face or enrolls a new one, keeping each person's ID stable over time.
gallery = []
next_id = 1
def recognize(embedding):
global next_id
best_index, best_sim = -1, 0.0
for index, (_, gallery_emb) in enumerate(gallery):
sim = float(np.dot(embedding, gallery_emb))
if sim > best_sim:
best_sim, best_index = sim, index
if best_sim >= MATCH_THRESHOLD:
person_id, gallery_emb = gallery[best_index]
# Blend the embedding into the gallery entry to stay robust to pose.
updated = 0.9 * gallery_emb + 0.1 * embedding
gallery[best_index] = (person_id, updated / np.linalg.norm(updated))
return person_id
person_id = next_id
next_id += 1
gallery.append((person_id, embedding))
print(f"Enrolled ID {person_id}", flush=True)
return person_id
def face_embeddings(video_frame):
"""Yield ((x, y, w, h), normalized_embedding) for each classified face."""
for region in video_frame.regions():
rect = region.rect()
emb = None
for tensor in region.tensors():
if tensor.is_detection():
continue
data = np.array(tensor.data(), dtype=np.float32)
if data.size >= 256:
emb = data[:256]
if emb is None:
continue
emb = emb / (np.linalg.norm(emb) + 1e-9)
yield (int(rect.x), int(rect.y), int(rect.w), int(rect.h)), emb
# Convert to BGR before inference so gvaclassify's embedding tensors survive to
# the appsink (a later format-changing videoconvert would strip them).
pipeline = Gst.parse_launch(
f"filesrc location={INPUT_VIDEO} ! decodebin3 ! "
"videoconvert ! video/x-raw,format=BGR ! "
f"gvadetect model={DETECTION_MODEL} device={DEVICE} "
f"threshold={DET_THRESHOLD} ! queue ! "
f"gvaclassify model={REID_MODEL} device={DEVICE} ! queue ! "
"appsink name=sink emit-signals=true sync=false max-buffers=4 drop=false"
)
sink = pipeline.get_by_name("sink")
writer = {"w": None}
def on_video(sink):
sample = sink.emit("pull-sample")
if sample is None:
return Gst.FlowReturn.OK
vf = VideoFrame(sample.get_buffer(), caps=sample.get_caps())
labeled = []
for (x, y, w, h), emb in face_embeddings(vf):
labeled.append((x, y, w, h, recognize(emb)))
with vf.data() as mat:
frame = mat.copy()
if writer["w"] is None:
frame_h, frame_w = frame.shape[:2]
structure = sample.get_caps().get_structure(0)
ok_fr, fps_n, fps_d = structure.get_fraction("framerate")
fps = fps_n / fps_d if ok_fr and fps_d else 12
writer["w"] = cv2.VideoWriter(
OUTPUT_VIDEO, cv2.VideoWriter_fourcc(*"mp4v"), fps, (frame_w, frame_h))
for x, y, w, h, person_id in labeled:
cv2.rectangle(frame, (x, y), (x + w, y + h), (0, 255, 0), 2)
cv2.putText(frame, f"ID {person_id}", (x, max(15, y - 8)),
cv2.FONT_HERSHEY_SIMPLEX, 0.6, (0, 255, 0), 2)
writer["w"].write(frame)
return Gst.FlowReturn.OK
sink.connect("new-sample", on_video)
pipeline.set_state(Gst.State.PLAYING)
pipeline.get_bus().timed_pop_filtered(
Gst.CLOCK_TIME_NONE, Gst.MessageType.EOS | Gst.MessageType.ERROR)
pipeline.set_state(Gst.State.NULL)
if writer["w"] is not None:
writer["w"].release()
print(f"Total identities recognized: {len(gallery)}", flush=True)
print(f"Saved: {OUTPUT_VIDEO}", flush=True)
Device targets:
DEVICE = "GPU"-- default in the sample code.DEVICE = "CPU"-- change"GPU"to"CPU".DEVICE = "NPU"-- change"GPU"to"NPU"; usebatch-size=1andnireq=4for best NPU utilization.
Expected Output
License
Licensed under the MIT License. See LICENSE for details.

