camera-motion-nano

47,238 parameters. 189 KB. What did the camera do β€” hold still, pan, zoom, or shake?

It never looks at a pixel. It reads the motion vectors the video encoder already computed and wrote into the bitstream.

Domain measured / deployment domain tested: measured on H.264 clips synthesised from COCO photographs and webcam frames with known affine motion; deployment domain: no real handheld, PTZ or vehicle-mounted camera motion tested. (Fifth line of the card standard, added 2026-09-02: a number is only as good as the domain it was measured in.)

The result worth reporting

input accuracy
motion vectors only (this model) 0.992
pixels β€” two 64Γ—64 frames, same architecture 0.785
best single scalar on the MV field, fitted in-sample 0.467
chance 0.167

Reading the codec's own metadata is not merely cheaper than looking at pixels here β€” it is more accurate, by +0.207. That is not surprising once stated: the encoder performed dense block-matching at full resolution as part of its ordinary work, while the pixel model sees two downsampled thumbnails. The expensive computation was already done and thrown away.

Per-class recall: static 1.000 Β· pan-x 0.984 Β· pan-y 1.000 Β· zoom-in 0.983 Β· zoom-out 1.000 Β· shake 0.984.

Scope

For: camera-tamper detection (has this fixed camera been moved?), footage triage and indexing, stabiliser gating, and activity sensing on hardware too weak to run a vision model.

Not for:

  • Not object or people tracking. It describes global frame motion, not what is in the scene.
  • Not a surveillance or identification tool. It sees a 16Γ—16 grid of displacements and cannot represent a person.
  • Not a substitute for IMU data where you have one.
  • It reports camera motion, never intent or content.

Input

A 16Γ—16Γ—2 motion-vector field, mean displacement per cell over the clip, resized to 64Γ—64 and fed as 2 channels (x, y). Extraction with PyAV:

import av, numpy as np, cv2, onnxruntime as ort
GRID, SZ = 16, 320
CLASSES = ["static","pan_x","pan_y","zoom_in","zoom_out","shake"]

def mv_field(path):
    acc = np.zeros((GRID,GRID,2), np.float32); cnt = np.zeros((GRID,GRID), np.float32)
    c = av.open(path); st = c.streams.video[0]
    st.codec_context.options = {"flags2": "+export_mvs"}       # REQUIRED
    for fr in c.decode(st):
        for sd in fr.side_data:
            if "Motion" not in type(sd).__name__: continue
            a = sd.to_ndarray()
            if not len(a): continue
            gx = np.clip((a["dst_x"]*GRID//SZ).astype(int), 0, GRID-1)
            gy = np.clip((a["dst_y"]*GRID//SZ).astype(int), 0, GRID-1)
            sc = np.maximum(a["motion_scale"], 1)
            np.add.at(acc,(gy,gx,0), a["motion_x"]/sc)
            np.add.at(acc,(gy,gx,1), a["motion_y"]/sc)
            np.add.at(cnt,(gy,gx), 1)
    f = acc/np.maximum(cnt,1)[...,None]
    return cv2.resize(f,(64,64),interpolation=cv2.INTER_NEAREST).transpose(2,0,1)

sess = ort.InferenceSession("camera_motion.onnx", providers=["CPUExecutionProvider"])
x = mv_field("clip.mp4")[None]
print(CLASSES[int(sess.run(None, {"input": x})[0][0].argmax())])

SZ must match your video's dimensions β€” the grid mapping divides by it. Vectors are scaled by motion_scale; skipping that silently changes the units.

Validated on a second sensor (2026-08-25)

Everything above used clips built from COCO photographs. Repeated on 320 clips built from Logitech BRIO frames of a real office β€” a different sensor, a different image pipeline, and imagery that had already been through the camera's own MJPG compression:

accuracy static pan_x pan_y zoom_in zoom_out shake
COCO photographs 0.990 1.00 1.00 1.00 1.00 1.00 0.94
Logitech BRIO 0.975 1.00 0.96 1.00 1.00 1.00 0.89

A gap of 0.015. That is a much smaller drop than sibling models in this family show on the same kind of test (the resolution model loses 0.106), and the reason is structural: this model never sees pixels. Motion vectors are the encoder's description of movement, so sensor noise, ISP sharpening and colour rendering are already abstracted away before the model looks at anything.

Be precise about what this does and does not establish. It shows the model transfers across sensors. It does not show transfer to real optical camera motion β€” the motion is still synthetic affine, as stated in the failure modes. Real handheld movement has rolling-shutter skew and non-rigid components that this has never seen.

Sensitivity floor β€” it measures VELOCITY, not drift

Added after an external validation that failed, which turned out to be informative rather than bad news. Measured on synthetic pans at known per-frame velocity, 30 clips per row:

px / frame total drift over the clip P(static) modal class
0.00 0.0 px 0.999 static
0.10 1.3 px 0.998 static
0.25 3.2 px 0.996 static
0.50 6.5 px 0.715 static
1.00 13.0 px 0.308 pan_x
2.00 26.0 px 0.115 pan_x
4.00 52.0 px 0.006 pan_x

Detection threshold is roughly 0.5–1.0 px per frame. Below 0.25 px/frame it reports static with high confidence, and it is not wrong to β€” motion vectors encode inter-frame displacement, and a fraction of a pixel per frame is genuinely no motion at the scale a codec works at.

It therefore CANNOT detect slow drift. Tested against footage whose cumulative camera drift had been measured independently by an ECC-based stabiliser at 61 px mean, this model called it static β€” correctly, because 61 px accumulated over roughly 82,000 frames is 0.0007 px/frame, about a thousand times below the floor above. The stabiliser measures displacement from a fixed reference; this model measures velocity between neighbours. They answer different questions and the test that conflated them was mine, not the model's.

If you need slow-drift detection, register frames against a fixed reference. This model is for motion that is happening now.

Known failure modes

  1. Trained on 14-frame clips at ~15 fps, H.264 CRF 23, 320Γ—320. Very different GOP structure, frame rate or resolution will shift the MV statistics. Re-check before trusting it elsewhere.
  2. Intra frames carry no motion vectors. A clip that is all I-frames yields an empty field. Confirmed on real footage: 11 of 12 frames carry MV side data, the I-frame correctly does not.
  3. Camera motion and large object motion are not separated. A close subject filling the frame and moving left produces a field resembling a pan.
  4. The motion is synthetic, the imagery and codec are not. Clips were built by applying known affine transforms to real COCO photographs and encoding them with a real H.264 encoder. Real handheld motion has rolling-shutter skew and non-rigid components this does not model.
  5. Six coarse classes. It will not give you a displacement in pixels.
  6. Slow drift is invisible β€” see the sensitivity floor above. This is the most likely way to misuse the model: a fixed camera creeping over hours reads as perfectly static.

How it was trained

  • 1,096 train / 368 test clips, split by source photograph so no image appears on both sides
  • Motion amplitude randomised per clip, so magnitude alone cannot answer the question
  • 4 conv layers (16β†’32β†’48β†’64), BatchNorm, global average pool, 6-way head
  • Adam 3e-3, 30 epochs, batch 64

The scalar rule was applied before publication: the best single-threshold classifier on the MV field scores 0.467 fitted in-sample. This model beats that optimistic baseline by +0.524 held out, which is why it exists rather than a threshold.

Verification

ONNX vs PyTorch, identical weights, both CPU, 256 inputs: max relative logit difference 2.1e-07, 100% argmax agreement.

Prior art β€” compressed-domain analysis is an established field

Using H.264 motion vectors instead of pixels is not a new idea. Compressed-domain video analysis has a substantial literature, including:

  • Moving-object detection and segmentation in the H.264/AVC compressed domain (Verstockt et al.; KΓ€s & Nicolas), operating by parsing syntax rather than decoding.
  • Compressed-domain human action recognition from motion vectors and quantisation parameters.
  • CoViAR (Wu et al., CVPR 2018), which trains directly on I-frames, motion vectors and residuals, and prior two-stream work replacing the optical-flow stream with a motion-vector stream precisely to avoid computing flow.

What is different here. The general finding that motion vectors substitute for optical flow is theirs. Compressed-domain work also targets content β€” human actions, moving objects β€” using substantially larger models. This targets global camera motion at a scale those papers do not work at, and reports two things they do not:

  • A controlled same-architecture comparison: 0.992 from motion vectors versus 0.785 from pixels, identical network, identical data, 47K parameters. The literature establishes that compressed-domain features work; this isolates how much comes from the representation alone.
  • A measured sensitivity floor: reliable below 0.25 px/frame, flipping near 1.0 px/frame. That bound is what tells you the model cannot see slow drift β€” it called footage with 61 px of cumulative drift static, correctly, because that is 0.0007 px/frame.

Practically, it is a 189 KB ONNX for camera-tamper and footage-triage, in a family where prior work ships papers and research code.

What "scalar baseline" means on this card

Every margin quoted here is against a stated baseline, because a margin without one is not a measurement. The baseline is the best single-threshold classifier over ten cheap statistics, fitted optimistically:

mean Β· std Β· lapvar Β· hf (high-frequency energy ratio) Β· grad (Sobel magnitude) Β· entropy Β· centre_edge Β· radial_slope Β· row_fft_peak Β· col_fft_peak

The last four are spatially aware, added after an earlier six-statistic baseline β€” all global aggregates β€” was found to systematically overstate model value on spatially structured tasks. A baseline that cannot see where anything is loses to a CNN by default. On one test task that flaw inflated an apparent margin from +0.060 to +0.261.

Two questions are asked with it, and they disagree:

  • in-sample β€” threshold fitted on the data it is scored on. Deliberately generous. Answers is there structure beyond a low-order statistic?
  • transferred β€” threshold fitted on the training corpus, applied unchanged to the target. Answers what should I ship? On one task the in-sample figure was 0.954 and the transferred figure 0.565.

Where this card quotes a single scalar figure without qualification, it is the in-sample one.

Provenance

Source imagery is COCO val2017 (public). Motion is synthesised. No personal data, no surveillance footage, and no recordings of identifiable people are involved.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support