File size: 1,493 Bytes
7bdbabf 66eabcc 960f294 66eabcc 7bdbabf 66eabcc 7bdbabf 66eabcc 7bdbabf 66eabcc | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 | ---
title: Data-use Extraction — WebGPU
emoji: 🐈
colorFrom: indigo
colorTo: blue
sdk: static
sdk_version: 1.0.0
pinned: false
license: apache-2.0
short_description: In-browser data-use extraction on WebGPU
---
# Data-use Extraction — WebGPU
Fully client-side demo that extracts data mentions from research text and
attributes each one with provenance and usage/impact — all running locally in
your browser on WebGPU. No server, no uploads; nothing leaves your machine.
## Pipeline
```
text
→ span model (GLiNER) → mention spans + specificity
→ LFM2.5-350M multitask → provenance {producer, year, geography, acronym}
+ usage/impact {data_type, usage_action, impact_label, usage_summary}
→ verbatim filter → render
```
## Models
| stage | model | HF repo |
|---|---|---|
| span detection | GLiNER | [`rafmacalaba/gliner_datause_extended-onnx`](https://huggingface.co/rafmacalaba/gliner_datause_extended-onnx) |
| attribution | LFM2.5-350M multitask | [`rafmacalaba/lfm2.5-350M-datause-multitask-onnx`](https://huggingface.co/rafmacalaba/lfm2.5-350M-datause-multitask-onnx) |
Models download from the Hub on first load and are cached in your browser's
IndexedDB. Requires a WebGPU-capable browser (Chrome/Edge with WebGPU enabled).
## Run locally
```bash
git clone https://github.com/rafmacalaba/ai4data-playground.git
cd ai4data-playground/web && python3 -m http.server 8000
# open http://localhost:8000
```
|