Faucon 1 🦅 (preview) — a scam screenshot detector that chooses where to look closer
Faucon ("falcon") scores screenshots for scams (fake crypto giveaways and presales, fake "free Robux / V-Bucks / Nitro" generators, fake investment platforms, fake wallet support pages…) — and, unlike a classic classifier, it decides by itself which region to zoom into before giving its final answer.
Red box: the region the model chose for its second look (a fake "free Nitro" gift and a look-alike link). Example generated with fictional names.
🧪 Preview / experimental. The learned zoom works (the model looks at text blocks, buttons, payment areas), but in this first version the second look changes the final score only slightly. At equal false-positive rate it is not yet better than our simpler production model Moineau. Use it to explore, not as your main filter.
How it "looks closer"
Everything happens inside the model (a single ONNX file, no extra logic needed):
- First look — the whole screenshot at normal resolution → a first opinion and an attention map ("how much do I want to look closer here?") over a 14 × 14 grid.
- Where — the attention map gives a center point (differentiable soft-argmax). Nobody tells the model where to look: it learned it only from being right or wrong.
- Zoom — a ×2 crop around that point is taken from the double-resolution input (real extra pixels).
- Second look — the same network looks at the zoomed crop; both looks are combined into the final score.
Quick start
pip install onnxruntime pillow numpy
python predict.py screenshot.png --show glimpses/ # saves images with the red zoom box
SCAM 0.999 zoom at (-0.39, -0.32) fake_nitro_dm.jpg
clean 0.696 zoom at (+0.05, +0.06) official_profile.png
Outputs of the ONNX model (input name image, float32 NCHW, dynamic batch / height / width):
| Output | Shape | Meaning |
|---|---|---|
logits |
(N, 2) | [clean, scam] → softmax, index 1 = scam probability |
glimpse_center |
(N, 2) | (x, y) in [-1, 1] of the zoomed region (box side = 1/2 of the image side) |
Preprocessing (see predict.py): RGB with transparency on white → longest side ≤ 768 px (bicubic) → closest
aspect-ratio bucket (height/width ∈ {0.5, 2/3, 1, 1.5, 2}) resized at double resolution (≈ 448 × 448 pixels, sides
multiple of 16, bilinear) → (x / 255 - 0.5) / 0.5.
Model
| Architecture | small CNN trained from scratch (own self-supervised pretraining on ~167 k unlabeled images, no external pretrained model) + learned attention / zoom + second look with shared weights |
| Parameters | ~1.17 M |
| Compute | ~2 × 0.5 GFLOP per image (two looks) |
| File | Faucon.onnx, 4.7 MB, opset 17 (GridSample inside the graph) |
| Recommended threshold | 0.85 (its scores are higher overall than Moineau's) |
Evaluation (real images, never seen during training)
| Test set | Threshold 0.85 (recommended) | Threshold 0.6 |
|---|---|---|
| 2,000 everyday clean images (art, web/app UIs, games, photos) | ≤ 9 false positives (≈ 0.45 %) | 29 false positives (1.5 %) |
| 36 real web scam screenshots | 11 / 36 caught | 21 / 36 caught |
| 30 legitimate X profiles | 0 / 30 false positives | 2 / 30 false positives |
| Discord moderation samples (70 clean images, 1 scam) | 0 false positives, scam caught | 1 false positive, scam caught |
Ranking quality on the 36 + 30 web set: AP 0.924 (Moineau: 0.907). The test sets are small: differences of 1–2 images are within noise.
Known limitations
- The second look has little influence yet (typical change: +0.00 to +0.03): it mostly confirms the first look.
- Higher false-positive rate than Moineau at the same threshold; flagged look-alikes include donation / sign-up forms with preset amounts, verified brand profiles on X, colorful mobile-game menus with currency counters.
- Small text: even with the ×2 zoom, very small print is hard to read.
- New scam styles may be missed; not designed for phishing pages identical to the real site (scam only in the URL), text-only messages or videos (for GIFs, score key frames and take the maximum).
Training data (summary)
- Self-supervised pretraining (no labels) on ~167,000 images: synthetic websites (WebSight), photos (COCO, mini-ImageNet), mobile app screens (RICO), gameplay frames, memes (Hateful Memes), desktop UI screenshots (WildGUI), plus the labeled set below without its labels.
- Supervised training on ~1,200 scam and ~14,800 clean images: real scam pages captured by public URL scanners (reviewed by hand), synthetic scam screenshots each with a legitimate twin, legitimate pages from official domains, varied clean images from public sources, and moderation samples from the author's own Discord bot.
The training data and training code are not released.
Intended use
- ✅ Research and experimentation on attention / "looking closer" for image moderation; flagging images for human review.
- ❌ Automatic punishment of users; commercial use; deciding on its own whether a website is safe.
License
CC BY-NC 4.0. Several pretraining and training sources only allow non-commercial / research use, hence this license.
Version
Faucon 1 — October 2026, first preview of the "thinking" line. Planned next steps: choosing the zoom level, several looks with memory, learning when to stop looking, and naming the clue it found ("fake airdrop", "game currency"…).
Citation
If you use Faucon in your work, please cite it (attribution is required by the CC BY-NC 4.0 license):
@misc{charlet2026faucon,
author = {Charlet, Th{\'e}o},
title = {Faucon: a scam screenshot detector that learns where to look closer},
year = {2026},
publisher = {Hugging Face},
howpublished = {\url{https://huggingface.co/RDTvlokip/Faucon}}
}
Théo CHARLET
TSSR Graduate (IT Systems & Networks Technician) — AI/ML Specialization
Creator of AG-BPE (Attention-Guided Byte-Pair Encoding)
🚀 Seeking internship opportunities
- Downloads last month
- 17
