Instructions to use litert-community/D-FINE-S-LiteRT with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- LiteRT
How to use litert-community/D-FINE-S-LiteRT with LiteRT:
# No code snippets available yet for this library. # To use this model, check the repository files and the library's documentation. # Want to help? PRs adding snippets are welcome at: # https://github.com/huggingface/huggingface.js
- Notebooks
- Google Colab
- Kaggle
D-FINE-S — LiteRT (CompiledModel GPU)
D-FINE (USTC, 2024 — ustc-community/dfine-small-coco), the SOTA
real-time DETR, converted to LiteRT and running 100% on the CompiledModel GPU (ML Drift) on a
phone, with no CPU/ONNX fallback.
D-FINE is a transformer detector — HGNetV2 backbone + a hybrid AIFI/CCFM encoder + an FDR (Fine-grained
Distribution Refinement) decoder. Off-the-shelf it is GPU-incompatible (deformable grid_sample →
GATHER_ND, two-stage query selection → TOPK/GATHER). Here it is converted with litert-torch and
split into two GPU graphs with a host step between them, so both transformer graphs run on the GPU.
Files
| File | What it is | Size (fp16) |
|---|---|---|
dfine_graphA_fp16.tflite |
HGNetV2 backbone + hybrid encoder + score head → enc_class[1,8400,80], memory_raw[1,8400,256] |
13.0 MB |
dfine_graphB_fp16.tflite |
two-stage combine + FDR decoder + heads → boxes[1,300,4] (cxcywh), logits[1,300,80] |
8.8 MB |
host_params.bin |
host per-token tail weights (enc_output + enc_bbox_head), valid mask, anchors (fp32) |
0.9 MB |
coco_labels.txt |
80 contiguous COCO class names (id 0–79) | — |
How it runs (two-graph split)
image[1,3,640,640]
→[GPU Graph A]→ enc_class, memory_raw
→[host: top-300 by max class score; per-token tail on the 300 selected (fp32):
target = enc_output(valid·memory_raw) (Linear + LayerNorm)
ref = enc_bbox_head(target) + anchors (3-layer MLP)]
→[GPU Graph B (memory_raw, target, ref)]→ boxes[1,300,4], logits[1,300,80]
→[host: sigmoid + threshold + cxcywh→xyxy + light NMS]→ detections
The on-device gate — a Mali 3D-sequence fan-out bug (NOT the FDR decoder)
A naïve Graph A (emitting enc_class/enc_coord/output_memory/memory_raw together) gave 0 detections
on device, and it first looked like the FDR decoder collapsing in fp16. That was a red herring. The real
cause is a Mali delegate bug: a 3-D token tensor [1,N,256] (from conv.flatten(2).transpose(1,2)) that
is both a graph output and consumed by another node — or that fans out to several consumers — is silently
clobbered on the longer branch (4-D conv-map outputs are fine). Here the raw memory output (Graph B's
cross-attention input) was garbage (device corr −0.02) → the decoder cross-attended to noise → no detections.
Fix: Graph A emits only the two fp16-clean leaves (enc_class + memory_raw×2) and the per-token tail
(enc_output + enc_bbox_head) runs on the host over the 300 selected tokens (exact, since per-token ops
commute with the gather). With clean memory the FDR decoder is perfect — correlation is not the ship
criterion, real-image detection IoU is.
Minimal usage
Android (Kotlin, CompiledModel GPU)
val ga = CompiledModel.create(context.assets, "dfine_graphA_fp16.tflite",
CompiledModel.Options(Accelerator.GPU), null)
val gb = CompiledModel.create(context.assets, "dfine_graphB_fp16.tflite",
CompiledModel.Options(Accelerator.GPU), null)
val aIn = ga.createInputBuffers(); val aOut = ga.createOutputBuffers()
val bIn = gb.createInputBuffers(); val bOut = gb.createOutputBuffers()
aIn[0].writeFloat(chw) // [1,3,640,640] RGB in [0,1], NCHW
ga.run(aIn, aOut) // -> enc_class[1,8400,80], memory_raw*2[1,8400,256]
// host step: /2 -> top-300 -> per-token tail (host_params.bin) -> target[1,300,256], ref[1,300,4]
// (resolve buffer slots by float size; full math in the Python below / litert-samples object_detection)
bIn[0].writeFloat(memory); bIn[1].writeFloat(target); bIn[2].writeFloat(ref)
gb.run(bIn, bOut)
val boxes = bOut[0].readFloat() // [1,300,4] cxcywh in [0,1]
val logits = bOut[1].readFloat() // [1,300,80] -> sigmoid + threshold + light NMS
Python (desktop verification)
import numpy as np
from PIL import Image
from ai_edge_litert.interpreter import Interpreter
NP_, NQ, NC, H = 8400, 300, 80, 256
img = Image.open("photo.jpg").convert("RGB").resize((640, 640))
x = (np.asarray(img, np.float32) / 255.0).transpose(2, 0, 1)[None] # [1,3,640,640], [0,1] only
# host_params.bin (fp32 LE): enc_output W[256,256],b,gamma,beta · bbox-MLP W0,b0,W1,b1,W2[4,256],b2 · valid[8400] · anchors[8400,4]
p = np.fromfile("host_params.bin", np.float32); o = 0
def take(*s):
global o; n = int(np.prod(s)); v = p[o:o+n].reshape(s); o += n; return v
eoW, eoB, eoG, eoBe = take(H, H), take(H), take(H), take(H)
W0, b0, W1, b1, W2, b2 = take(H, H), take(H), take(H, H), take(H), take(4, H), take(4)
valid, anchors = take(NP_), take(NP_, 4)
def run(path, feeds): # feed/fetch tensors by shape (converter slot order is arbitrary)
it = Interpreter(model_path=path); it.allocate_tensors()
for d in it.get_input_details(): it.set_tensor(d["index"], feeds[tuple(d["shape"][1:])])
it.invoke(); return {tuple(d["shape"][1:]): it.get_tensor(d["index"]) for d in it.get_output_details()}
a = run("dfine_graphA_fp16.tflite", {(3, 640, 640): x})
enc_cls, mem = a[(NP_, NC)][0], a[(NP_, H)][0] / 2.0 # Graph A emits memory_raw*2 — undo
top = np.argsort(-enc_cls.max(-1))[:NQ] # top-300 by max class logit
t = (valid[top, None] * mem[top]) @ eoW.T + eoB # per-token tail: enc_output Linear...
t = (t - t.mean(-1, keepdims=True)) / np.sqrt(t.var(-1, keepdims=True) + 1e-5) * eoG + eoBe # ...+ LayerNorm
h = np.maximum(t @ W0.T + b0, 0); h = np.maximum(h @ W1.T + b1, 0)
ref = h @ W2.T + b2 + anchors[top] # enc_bbox_head MLP + anchors
b = run("dfine_graphB_fp16.tflite",
{(NP_, H): mem[None], (NQ, H): t[None].astype(np.float32), (NQ, 4): ref[None].astype(np.float32)})
boxes, logits = b[(NQ, 4)][0], b[(NQ, NC)][0] # cxcywh in [0,1] / 80-way logits
labels = open("coco_labels.txt").read().splitlines()
score = 1 / (1 + np.exp(-logits.max(-1))); cls = logits.argmax(-1)
for q in np.where(score > 0.4)[0]: # + light NMS (IoU 0.7) in a real app
cx, cy, w, hh = boxes[q]
print(f"{labels[cls[q]]:12s} {score[q]:.2f} xyxy=({cx-w/2:.3f},{cy-hh/2:.3f},{cx+w/2:.3f},{cy+hh/2:.3f})")
On-device (Pixel 8a, Tensor G3 — verified)
Both graphs run 100% GPU-resident (LITERT_CL): Graph A 511/511, Graph B 850/850. On a COCO val image
(giraffe + cows) the device chain reproduces the PyTorch detections at IoU 0.99–1.00 with matching class
and score. End-to-end ~450 ms/frame — accurate and fully-GPU but not real-time on this device (the
deformable decoder over the 8400 tokens / 80×80 levels is GPU-compute-bound; the GATHER-free tent-matmul
grid_sample turns an O(points) gather into an O(H·W) matmul). For a real-time camera DETR see
RF-DETR Nano.
Preprocessing / outputs
- Input: square resize to 640×640, RGB,
[0,1]rescale only (no ImageNet normalization), NCHW. - Output: Graph B
boxesarecxcywhnormalized to[0,1];logitsare 80-way (contiguous COCO id 0–79). Host applies sigmoid + score threshold +cxcywh→xyxy+ light NMS.
Conversion notes
Converted with litert-torch (NCHW preserved — onnx2tf destroys ViT attention). Re-authoring (per-graph
tflite-vs-torch correlation 1.0): deformable grid_sample → a GATHER/CAST-free tent-matmul, multi-level
MSDeformAttn ≤4D, the FDR LQE prob.topk → iterative max-and-mask, distance2bbox stack→cat, baked AIFI
sine pos-embed, a down-scaled fp16-safe LayerNorm, and the 3D-fan-out fix above (emit clean leaves +
host-side per-token tail).
A runnable Android sample (CompiledModel GPU) and the conversion scripts are in the official
ai-edge-litert/litert-samples object_detection example.
Performance
Measured on a Pixel 8a (Tensor G3, Android 16) with the standard TFLite benchmark_model tool — 10 warm-up runs then 50 timed runs, reported as the tool's mean.
| Runtime | Backend | Graph on GPU | Latency |
|---|---|---|---|
TFLite benchmark_model (TfLiteGpuDelegateV2) — dfine_graphA_fp16.tflite |
GPU (OpenCL) | 233 / 511 | 2193.2 ms |
TFLite benchmark_model (TfLiteGpuDelegateV2) — dfine_graphB_fp16.tflite |
GPU (OpenCL) | 128 / 850 | did not run |
TFLite benchmark_model — dfine_graphA_fp16.tflite |
CPU (XNNPACK, 4 threads) | — | XNNPACK declined the graph |
TFLite benchmark_model — dfine_graphB_fp16.tflite |
CPU (XNNPACK, 4 threads) | — | 321.8 ms |
Any on-device figure recorded when this model shipped came from a different runtime. It was taken through LiteRT's own CompiledModel accelerator (logcat reports it as LITERT_CL), which is the path the Kotlin sample app and the LiteRT API use, and it appears elsewhere on this card. The rows above are the classic TFLite OpenCL delegate, measured with a tool anyone can download and re-run. The two are not comparable, so read the rows above as a reproducible floor rather than as this model's speed on LiteRT.
XNNPACK declines these fp16 graphs — it reports failed to delegate DEPTHWISE_CONV_2D and then fails to allocate tensors — so there is no usable CPU number. Disabling XNNPACK falls back to reference kernels, which measured about 20× slower than the GPU on models of this size and would not represent CPU inference anyone would ship.
Note that the GPU does not take the whole graph here (233 / 511 in dfine_graphA_fp16.tflite, 128 / 850 in dfine_graphB_fp16.tflite); the remainder runs on the CPU and the split costs a per-partition round trip.
License
Apache-2.0, inherited from Peterande/D-FINE.
- Downloads last month
- 63
Model tree for litert-community/D-FINE-S-LiteRT
Base model
ustc-community/dfine-small-coco