Faucon 1 🦅 (preview) — a scam screenshot detector that chooses where to look closer

Faucon ("falcon") scores screenshots for scams (fake crypto giveaways and presales, fake "free Robux / V-Bucks / Nitro" generators, fake investment platforms, fake wallet support pages…) — and, unlike a classic classifier, it decides by itself which region to zoom into before giving its final answer.

Faucon zooms on the fake Nitro gift and the look-alike link

Red box: the region the model chose for its second look (a fake "free Nitro" gift and a look-alike link). Example generated with fictional names.

🧪 Preview / experimental. The learned zoom works (the model looks at text blocks, buttons, payment areas), but in this first version the second look changes the final score only slightly. At equal false-positive rate it is not yet better than our simpler production model Moineau. Use it to explore, not as your main filter.

How it "looks closer"

Everything happens inside the model (a single ONNX file, no extra logic needed):

  1. First look — the whole screenshot at normal resolution → a first opinion and an attention map ("how much do I want to look closer here?") over a 14 × 14 grid.
  2. Where — the attention map gives a center point (differentiable soft-argmax). Nobody tells the model where to look: it learned it only from being right or wrong.
  3. Zoom — a ×2 crop around that point is taken from the double-resolution input (real extra pixels).
  4. Second look — the same network looks at the zoomed crop; both looks are combined into the final score.

Quick start

pip install onnxruntime pillow numpy
python predict.py screenshot.png --show glimpses/      # saves images with the red zoom box
SCAM   0.999  zoom at (-0.39, -0.32)  fake_nitro_dm.jpg
clean  0.696  zoom at (+0.05, +0.06)  official_profile.png

Outputs of the ONNX model (input name image, float32 NCHW, dynamic batch / height / width):

Output Shape Meaning
logits (N, 2) [clean, scam] → softmax, index 1 = scam probability
glimpse_center (N, 2) (x, y) in [-1, 1] of the zoomed region (box side = 1/2 of the image side)

Preprocessing (see predict.py): RGB with transparency on white → longest side ≤ 768 px (bicubic) → closest aspect-ratio bucket (height/width ∈ {0.5, 2/3, 1, 1.5, 2}) resized at double resolution (≈ 448 × 448 pixels, sides multiple of 16, bilinear) → (x / 255 - 0.5) / 0.5.

Model

Architecture small CNN trained from scratch (own self-supervised pretraining on ~167 k unlabeled images, no external pretrained model) + learned attention / zoom + second look with shared weights
Parameters ~1.17 M
Compute ~2 × 0.5 GFLOP per image (two looks)
File Faucon.onnx, 4.7 MB, opset 17 (GridSample inside the graph)
Recommended threshold 0.85 (its scores are higher overall than Moineau's)

Evaluation (real images, never seen during training)

Test set Threshold 0.85 (recommended) Threshold 0.6
2,000 everyday clean images (art, web/app UIs, games, photos) ≤ 9 false positives (≈ 0.45 %) 29 false positives (1.5 %)
36 real web scam screenshots 11 / 36 caught 21 / 36 caught
30 legitimate X profiles 0 / 30 false positives 2 / 30 false positives
Discord moderation samples (70 clean images, 1 scam) 0 false positives, scam caught 1 false positive, scam caught

Ranking quality on the 36 + 30 web set: AP 0.924 (Moineau: 0.907). The test sets are small: differences of 1–2 images are within noise.

Known limitations

  • The second look has little influence yet (typical change: +0.00 to +0.03): it mostly confirms the first look.
  • Higher false-positive rate than Moineau at the same threshold; flagged look-alikes include donation / sign-up forms with preset amounts, verified brand profiles on X, colorful mobile-game menus with currency counters.
  • Small text: even with the ×2 zoom, very small print is hard to read.
  • New scam styles may be missed; not designed for phishing pages identical to the real site (scam only in the URL), text-only messages or videos (for GIFs, score key frames and take the maximum).

Training data (summary)

  • Self-supervised pretraining (no labels) on ~167,000 images: synthetic websites (WebSight), photos (COCO, mini-ImageNet), mobile app screens (RICO), gameplay frames, memes (Hateful Memes), desktop UI screenshots (WildGUI), plus the labeled set below without its labels.
  • Supervised training on ~1,200 scam and ~14,800 clean images: real scam pages captured by public URL scanners (reviewed by hand), synthetic scam screenshots each with a legitimate twin, legitimate pages from official domains, varied clean images from public sources, and moderation samples from the author's own Discord bot.

The training data and training code are not released.

Intended use

  • ✅ Research and experimentation on attention / "looking closer" for image moderation; flagging images for human review.
  • ❌ Automatic punishment of users; commercial use; deciding on its own whether a website is safe.

License

CC BY-NC 4.0. Several pretraining and training sources only allow non-commercial / research use, hence this license.

Version

Faucon 1 — October 2026, first preview of the "thinking" line. Planned next steps: choosing the zoom level, several looks with memory, learning when to stop looking, and naming the clue it found ("fake airdrop", "game currency"…).

Citation

If you use Faucon in your work, please cite it (attribution is required by the CC BY-NC 4.0 license):

@misc{charlet2026faucon,
  author       = {Charlet, Th{\'e}o},
  title        = {Faucon: a scam screenshot detector that learns where to look closer},
  year         = {2026},
  publisher    = {Hugging Face},
  howpublished = {\url{https://huggingface.co/RDTvlokip/Faucon}}
}

Théo CHARLET

TSSR Graduate (IT Systems & Networks Technician) — AI/ML Specialization
Creator of AG-BPE (Attention-Guided Byte-Pair Encoding)

LinkedIn Website Search

🚀 Seeking internship opportunities

Downloads last month
17
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support