litert-models / docs /CASCADE_V25.md
unicornwhodev's picture
Add four verified v2.5 LiteRT cascade models, Cadryl contracts and bilingual 680-image benchmark
17133be verified
|
Raw History Blame Contribute Delete
15 kB

FireViewer v2.5 — Cascade / Cascade complète

English

Four frozen, selected checkpoints are now exported to inference-only LiteRT FP32. The source PyTorch revisions and safetensor digests are recorded per package. No new training, quantization, calibration or threshold tuning accompanies the conversion. Ordinary consumers use runtime_contract.json, per-output label maps and preprocessor_config.json; Cadryl uses its exact android_model_config.json schema or studio_pack.json (vision-studio-pack/1). The six historical packages remain available with their original scope and evidence.

Stage / Étape LiteRT package Input CPU numeric parity Android device
proposal fireviewer_rfdetr_medium_probable_fire_v25_fp32 1088 px · batch 1 PASS at .10 operating threshold · 11 inputs Not qualified / Non qualifié
visible fireviewer_rfdetr_medium_visible_flame_v25_fp32 1088 px · batch 1 PASS at .25 operating threshold · 11 inputs Not qualified / Non qualifié
point fireviewer_dinov2_small_visible_flame_monopoint_v25_fp32 448 px · batch 1 PASS tensor outputs · 5 inputs Not qualified / Non qualifié
smoke fireviewer_dinov2_small_fire_associated_smoke_v25_fp32 448 px · batch 1 PASS tensor outputs · 11 inputs Not qualified / Non qualifié

Functional routing

flowchart TD
    I[RGB whole image] --> P[RF-DETR M probable_fire · 1088]
    P --> R[score .10 · NMS .75 · at most 6 ROI]
    R --> V[30% context · min 128px · RF-DETR M visible_flame · 1088]
    V --> F[Flame boxes · score .25 · reprojection · global NMS .60]
    F --> K[Tight flame crops · DINOv2-S heatmap · 448]
    K --> A[One flame_base_anchor per detected flame]
    V --> G[ROI without a local confirmed flame]
    G --> S[45% context · min 160px · DINOv2-S smoke classifier · 448]
    S --> T[Calibrated compatibility · low / ambiguous / high]

Context is added on each side: expanded width and height are multiplied by 1 + 2 × context, before the minimum side and image clipping. RF-DETR receives NCHW float32 RGB resized with half-pixel bilinear interpolation, no antialias, with RGB converted to float before resize; normalization is ImageNet mean [.485,.456,.406], std [.229,.224,.225]. Point crops use Pillow bilinear; smoke crops use Pillow bicubic. All models stretch to the specified square size. The supplied NumPy/Pillow runner preserves these differences.

Proposal score ≥0.10, NMS IoU 0.75, limit 6; visible score ≥0.25 and global image-space NMS 0.60, limit 100. Tight point crops use floor for the minimum and ceil for the maximum. The heatmap point is the argmax followed by a local 5×5 softmax expectation at heatmap pixel centers. It is an appearance anchor, not an ignition location, geographic position or independently calibrated source confidence.

Smoke temperature is 1.1490650177001953. Low compatibility ≤0.30160276398937463; high compatibility ≥0.8199750333466715; values between these are ambiguous and do not raise the benchmark image alarm. Smoke is a binary visual compatibility classifier. It supplies no mask, smoke-core, smoke point or origin. No learned fire-group association is performed.

Conventional format and Cadryl

Every package contains model.tflite, actual input/output shapes, output positions, class maps, preprocessing, source revision/digests, validation and a SHA256 manifest. The conventional detector outputs are all 300 normalized cxcywh boxes and raw class logits; active class is index 0, index 1 is unused background. The supplied runner reproduces the stable native global top-300 query/class sort.

Cadryl's RF-DETR adapter skips index 0. The compatible graph therefore also exposes YOLO-style BCN [1,5,300] (cx,cy,w,h,active probability), preserving RF-DETR weights and architecture. The graph selects the global top 300 sigmoid query/class scores and zeros scores for unused background. Cadryl applies its configured threshold/NMS. For the visible specialist this is single-crop annotation; the complete cascade must reproject first and apply global NMS through the runner. Importing one model in Cadryl does not automatically orchestrate all four models.

Detector conversion qualification is scoped to the frozen operating points (.10 / .25). All-300 raw-query numerical parity fails for a few very-low-score candidates, even after set correspondence. These changes are consistent with near-tied two-stage top-k selection across FP32 backends (a diagnostic inference). Operative detection counts, boxes and scores pass the original numerical bounds; low-score raw discrepancies remain disclosed in each conversion_validation.json and raw-query-diagnostic.json. Query row identity and lower score cutoffs are not qualified. The DINOv2 heatmap/point and calibrated probabilities pass tensor-output comparisons. No numerical tolerance was loosened to hide raw-query changes.

The point graph exposes [1,1,2] normalized x/y and raw [1,1,128,128] heatmap logits. A two-coordinate Cadryl point has an implicit score of 1; this is structural presence, not calibrated confidence. Use the point model only on tight confirmed flame crops. The smoke graph exposes raw logit, complementary calibrated probabilities, and [1,1] positive probability. Cadryl annotates only high positive compatibility, while host code retains low/ambiguous/high triage.

Cadryl Android scales bitmaps with bilinear integer-pixel interpolation. Exact float-resize detector parity and bicubic smoke pixel parity on Android are not demonstrated. CPU LiteRT invocation and ModelConfig compatibility do not qualify an Android phone, APK, GPU/NPU delegate, memory budget or application accuracy. These four new graphs have no train/save/restore signatures. Do not apply the historical trainable-head SDK contract to them.

Reproduce inference

pip install ai-edge-litert==2.2.0 numpy pillow huggingface_hub
hf download fireviewer/litert-models --revision <publication-revision> --local-dir ./fireviewer-litert
python fireviewer-litert/cascade/v2.5/cascade_litert.py --root ./fireviewer-litert --image image.jpg --output predictions.json --threads 2

Pin the immutable publication revision from reports/cascade-v25-litert-publication.json. Check each package manifest before inference. Loading the models and each model's numerical conversion tests are separate from semantic evaluation. No test split was used during conversion. Conversion probes include annotation-derived crops solely for tensor regression; they are not a proposals-only semantic validation set.

The complete LiteRT runner is also compared with eager PyTorch on the same CPU on six preselected images (two distinct visual families per stratum), exercising all four stages. The original L4 outputs are retained as a separate cross-backend diagnostic. Small detector differences can cross an integer crop boundary and amplify downstream probability changes; exact end-to-end L4 equivalence is not presumed. These six numerical-regression cases do not replace the original 680-image semantic benchmark or constitute a full LiteRT rebenchmark. See routing evidence.

Full-cascade benchmark — 680 images

Frozen PyTorch FP32 pipeline on L4; no fitting or threshold selection on this panel. Case mix: 179 flame, 70 smoke, 431 negative images, 229 visual families. There may be several views of a family; families are not verified independent fire incidents.

Endpoint Result Numerator / denominator
Positive-image alarm recall 88.35% 220 / 249
Sampled image alarm precision 81.78% 220 / 269
Negative-image false-alarm rate 11.37% 49 / 431
Flame-image any-branch alarm recall 99.44% 178 / 179
Smoke-only any-branch alarm recall 60.00% 42 / 70
Flame bbox recall, IoU ≥ .5 25.64% 231 / 901
Flame bbox precision, IoU ≥ .5 27.57% 231 / 838
Tiny flame recall 22.59% 178 / 788
Multi-flame instance recall 24.72% 219 / 886
Expanded-proposal containment 97.89% 882 / 901
Point inside matched GT flame bbox 97.84% 226 / 231
Smoke-only image with false flame 10.00% 7 / 70

An image alarm is any visible-flame output or high smoke compatibility; it is not a statement that every flame has a precise box. There were TP=220, FP=49, FN=29, TN=382 for image alarms; flame localization TP=231, FP=607, FN=670. Tiny means bbox area/image area ≤ (32/1088)^2, a protocol bin rather than original-pixel COCO-small. Proposal containment means at least 90% of the GT box lies within an expanded proposal. The 229-image one-representative-per-family panel still has tiny recall 101/399=25.31%.

Interpretation: stage 1 generally provides coverage, while precise instance separation/localization of small visible flames remains the primary weakness. In the inspected worst-error gallery, several small annotated flames are merged into a broader detection. This is an observed error pattern, not a proven causal attribution for every miss. Good image suspicion and low box recall can coexist. The point-in-box statistic measures consistency only: no independent human flame-base GT is available, so PCK and base accuracy are unmeasured. Smoke-only labels establish visible smoke, not independently verified fire association.

Mean end-to-end inference latency was 226.82 ms/image, median 172.44 ms, P95 458.01 ms on L4 PyTorch. All invoked stages, decoding, preprocessing, NMS and reprojection are included; model loading/warmup are excluded. Stage calls: proposal 680, visible 1636, point 838, smoke 855. These are not LiteRT or Android performance figures. Separate CPU LiteRT numerical-test timings are recorded per conversion, with different samples and workload.

The image-level SHA/perceptual novelty audit covered all known historical splits including private sources and legacy FASDD/Pyro-SDIS manifests. Detected neighbors and suspected FLAME1 scenes were excluded before inference. No detected duplicate proves neither unseen events nor complete unknown provenance. Sources: DIFD and SAINetset v8, whose images may include synthetic/donated material. Only 208 retained images had a visual QA pass; there is no independent human annotation audit. Ratios describe this selected correlated panel, with no independent-event confidence intervals or deployment certification.

Full report · Machine-readable metrics · Frozen release manifest

Descriptive image-alert and flame-localization metrics

Français

La cascade utilise quatre checkpoints figés : propositions probable_fire sur image entière, spécialiste visible_flame dans les ROI, ancre flamme sur crop serré, puis classifieur fumée uniquement pour les ROI sans flamme locale confirmée. Les boîtes flamme sont reprojetées et dédoublonnées dans l'image entière. La fumée fournit un état bas/ambigu/haut ; aucun masque ni point fumée rejeté n'est publié. Les seuils restent ceux du benchmark, sans nouvel entraînement ni ajustement sur le lot.

Chaque export FP32 dispose d'un contrat conventionnel, des classes, des prétraitements exacts, d'un reçu de validation CPU et de SHA256. Les fichiers android_model_config.json et studio_pack.json utilisent les schémas réels de Cadryl. Le tensor BCN compatible YOLO ne change pas l'architecture RF-DETR : il évite que son indice de classe active 0 soit ignoré par l'adaptateur RF-DETR actuel. Le point normalisé est une ancre d'apparence de flamme, pas une origine d'incendie. Son score implicite Cadryl n'est pas une confiance calibrée. La classification fumée conserve sa température et ses deux seuils ; Cadryl ne propose que les positifs hauts.

La comparaison numérique LiteRT/PyTorch confirme le calcul sur les entrées testées. Elle ne démontre ni la précision des annotations ni le fonctionnement sur téléphone. Les différences de redimensionnement Android, les délégués GPU/NPU, la mémoire et la latence Android restent à qualifier. Les quatre nouveaux exports sont réservés à l'inférence, sans signatures d'apprentissage. L'import d'un modèle individuel dans Cadryl n'exécute pas automatiquement la cascade entière.

La qualification RF-DETR concerne les détections aux seuils figés 0,10 / 0,25. Quelques candidats bruts de très faible score diffèrent entre les backends ; l'équivalence numérique des 300 candidats n'est donc pas validée. Ces écarts et les seuils non qualifiés sont conservés dans les reçus. Les sorties point/heatmap et les probabilités DINOv2 passent les comparaisons de tenseurs. Aucun seuil ni tolérance n'a été modifié pour masquer les écarts.

Sur 680 images, l'alerte récupère 88,35% des images positives mais produit une alerte sur 11,37% des négatifs. La localisation à IoU ≥ 0,5 retrouve seulement 25,64% des boîtes flamme et 22,59% des petites flammes. Les propositions couvrent pourtant 97,89% des boîtes GT : le principal besoin est donc la séparation/localisation précise des instances par le spécialiste visible, sous cette distribution difficile. Les cas inspectés montrent notamment des petites flammes réunies dans une boîte trop large.

Les 97,84% de points dans une boîte flamme GT appariée sont un contrôle de cohérence, pas une mesure de justesse de base du feu : les points GT humains indépendants manquent. Les 229 familles visuelles, les vues corrélées, les annotations de source et les données SAI potentiellement synthétiques limitent la portée de ce diagnostic. Le test historique reste verrouillé. La latence médiane 172,44 ms/image concerne PyTorch sur L4, pas LiteRT ou Android.

Sources / Références

The original cards contain training recipes, selected-validation results, publication-subset differences and checkpoint provenance. Dataset rights remain source-specific. RF-DETR and DINOv2 source notices and Apache-2.0 licences are retained per package.

LiteRT PyTorch conversion · Official litert-torch · RF-DETR · DINOv2