|
Download README.md from rafmacalaba/datause-extraction-webgpu: direct link, hf CLI and curl.
- Browser
- Download file 1.49 kB
-
https://huggingface.co/spaces/rafmacalaba/datause-extraction-webgpu/resolve/main/README.md
- Command line
-
hf download hf://spaces/rafmacalaba/datause-extraction-webgpu/README.md
-
curl -L -o README.md https://huggingface.co/spaces/rafmacalaba/datause-extraction-webgpu/resolve/main/README.md
1.49 kB
| title: Data-use Extraction — WebGPU | |
| emoji: 🐈 | |
| colorFrom: indigo | |
| colorTo: blue | |
| sdk: static | |
| sdk_version: 1.0.0 | |
| pinned: false | |
| license: apache-2.0 | |
| short_description: In-browser data-use extraction on WebGPU | |
| # Data-use Extraction — WebGPU | |
| Fully client-side demo that extracts data mentions from research text and | |
| attributes each one with provenance and usage/impact — all running locally in | |
| your browser on WebGPU. No server, no uploads; nothing leaves your machine. | |
| ## Pipeline | |
| ``` | |
| text | |
| → span model (GLiNER) → mention spans + specificity | |
| → LFM2.5-350M multitask → provenance {producer, year, geography, acronym} | |
| + usage/impact {data_type, usage_action, impact_label, usage_summary} | |
| → verbatim filter → render | |
| ``` | |
| ## Models | |
| | stage | model | HF repo | | |
| |---|---|---| | |
| | span detection | GLiNER | [`rafmacalaba/gliner_datause_extended-onnx`](https://huggingface.co/rafmacalaba/gliner_datause_extended-onnx) | | |
| | attribution | LFM2.5-350M multitask | [`rafmacalaba/lfm2.5-350M-datause-multitask-onnx`](https://huggingface.co/rafmacalaba/lfm2.5-350M-datause-multitask-onnx) | | |
| Models download from the Hub on first load and are cached in your browser's | |
| IndexedDB. Requires a WebGPU-capable browser (Chrome/Edge with WebGPU enabled). | |
| ## Run locally | |
| ```bash | |
| git clone https://github.com/rafmacalaba/ai4data-playground.git | |
| cd ai4data-playground/web && python3 -m http.server 8000 | |
| # open http://localhost:8000 | |
| ``` | |