File size: 6,714 Bytes
086a318
 
 
 
 
 
 
 
 
53ea208
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
676999c
 
f510dfc
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
---
title: Data-use annotation review
emoji: 🏷️
colorFrom: blue
colorTo: green
sdk: static
pinned: false
---

# human_labeling

Fork of `analysis/v24_sample_review/audit_deck.html` as a hostable
annotation pipeline. Same blind KEEP/DROP interaction, but the queue
comes from the Hub instead of an 11MB embedded `ROWS` array.

- Source: `rafmacalaba/datause-ner` ([viewer](https://huggingface.co/datasets/rafmacalaba/datause-ner/viewer/probe_splits))
  for passages; **scores are singlepass only** — the infer head in
  `rafmacalaba/gliner-datause-catchall-singlepass`
  (`sigmoid(head([start; end; mean; ±64 window]))` on inference-native features).
  No v3/encode-head scores anywhere in this tool.
- Default queue: `probe_candidates/pool` — the only config with an
  embedded `passage`. `probe_splits` train/val/holdout rows have **no
  passage** (surface + word offsets only), so they annotate surface-only
  unless joined to a chunk corpus (see caveat below).
- Annotations land in `human_labeling/annotations/rulings.jsonl`
  (one JSON object per ruling). On static hosting (github.io) there is
  no writer: the app exports a JSON file the annotator drops into that
  folder (PR). On Spaces the bundled `app.py` accepts `POST
  /api/rulings` and appends server-side.

## Quick start (local)

```bash
uv run python human_labeling/build_gliner_queue.py            # Luna spans + singlepass scores
uv run python human_labeling/app.py  # serves index.html + /api/queue + POST /api/rulings
```

Data source: the `gliner` config of rafmacalaba/datause-ner (tokenized
passages + catch-all entity spans with camp2 Luna traceability). The queue
maps each Luna verdict onto its span, rescores with the singlepass head
(probe score) and the GLiNER proposer (extractor score @0.1), and the UI
reviews keep/drop with both signals. Retired experiments (pool, passage
grouping, gliner2/human473 queues) live only in git history.

## Hosting

- **github.io (static):** publish `human_labeling/` as-is. `index.html`
  fetches `./queue_gliner.json` (built by `build_gliner_queue.py`, committed). Rulings
  persist in `localStorage`; Export downloads
  `rulings-<annotator>.json`, which is committed under `annotations/`
  and reconciled with `merge.py`.
- **Hugging Face Spaces (SDK: Docker, port 7860):** same folder plus
  `app.py`. The frontend feature-detects `POST /api/rulings`: if it
  200s, rulings stream to `annotations/rulings.jsonl` on disk; if not,
  it falls back to export. Dockerfile provided.

## Files

| file | purpose |
|---|---|
| `index.html` | annotation UI (passage view, per-mention ruling, luna×model filter) |
| `build_gliner_queue.py` | builds + scores `queue_gliner.json` from the gliner config (singlepass probe + extractor on MPS) |
| `queue_gliner.json` | served queue: 225 passages / 522 Luna-labeled spans, round-robin over origins, multi-mention first |
| `probe_labels.py` | shared decision rule + per-origin singlepass thresholds |
| `app.py` | stdlib server: static + `GET /api/queue` + `POST /api/rulings` |
| `merge.py` | `annotations/rulings*.jsonl` → `annotations/adjudicated.jsonl` (majority per key) |
| `annotations/` | ruling storage; every file is JSONL, one ruling per line |
| `Dockerfile` | Spaces deploy (Docker SDK) |
| `build_gliner2_queue.py`, `build_pool_queue.py`, `rescore_singlepass.py`, `build_passage_queue.py`, `build_gold_queue.py` | retired experiments (kept for reference; not served) |

## Caveat: `probe_splits` has no context

`probe_train/val/holdout.jsonl` rows are
`key, origin, corpus_id, w_start, w_end, surface, label, source,
head_score, split` — no passage text. The underlying chunk text lives
in per-corpus repos (`rafmacalaba/fcv-extractions-meta[-tiered]`), and
`general_prwp` (the majority origin) has no published chunk source at
all (same gap as `demo/server.py: HOLDOUT_META_SOURCES`). So:

- `build_pool_queue.py --config probe_candidates` (default): full passage, blind ruling quality = audit_deck.
- `build_pool_queue.py --config probe_splits`: surface-only triage; items are
  flagged `"ctx_missing": true` and the UI shows the surface + scores
  without a context card. Fine for obvious keeps/drops, not for adjudication.

## Bands (decision labels)

Every queue item carries `band`, tagged from the singlepass score with the
bundle's per-origin best-F1 thresholds (`probe_labels.py`):

| band | rule |
|---|---|
| `keep` | `head_score >= per-origin threshold` — model-flagged keep, quick confirm |
| `confusion` | `DROP_FLOOR < head_score < threshold` — the review zone (was "mid") |
| `drop` | `head_score <= 0.05` — confident reject, quick confirm |
| `unscored` | score missing (out-of-grid span) |

Per-origin thresholds (published `thresholds.json`): `fcv_pads_east_africa`
0.5 · `general_prwp` 0.4 · `jad_paddy_docs` 0.1 · `jdc_operational` 0.6 ·
`refugee_pads` 0.5 · `reliefweb` 0.4; fallback `global_best` 0.5.

## Ruling schema

```json
{"key": "sample:fcv_pads_east_africa:002082:7:1:0", "ruling": "DATA_MENTION",
 "annotator": "rm", "ts": "2026-09-08T12:00:00Z", "note": "",
 "head_score": 0.61, "origin": "fcv_pads_east_africa", "queue": "probe_candidates/pool"}
`ruling` is `DATA_MENTION` (keep) or `NON_MENTION` (drop) — same labels as the deck export.
## review.html (current)

Multi-span review over `rafmacalaba/data-use-ner`: `gliner` holdout (gold 473),
`to_annotate` pool, per-user `annotate_{paddy,aj,aivin,rafael}[_part2]` sets
(240 passages each, mutually exclusive, stratified by origin × probe tag).
Assignment manifests live in `assignment_manifest.json` (verified equal to
the Hub configs — queue files are append-only; never resampled in place).

### For annotators

1. Open your link (or pick your name on the landing page) — your file loads
   automatically. Type your name if asked; it locks on first ruling.
2. Judge each highlighted mention in its window: `→` next mention,
   `k` keep, `d` drop, `space` skip, `[`/`]` slide the window,
   `-`/`+` resize it. Optional freeform note per ruling.
3. Before ruling you see only `probe: easy` or `probe: confusion`. After
   ruling, Luna's verdict + reasoning appear, with a DISAGREE flag (also a
   toast) when you differ.
4. Every 10% you'll be nudged to **export** a checkpoint (`e` or the button,
   header and footer). Send `rulings-<name>-<queue>.jsonl` to rafael/aivin.
   To resume later, **import** the file back — you jump to the first
   unreviewed mention.

### For coordinators

Drop returned files into `human_labeling/annotations/` and run `merge.py` →
majority verdict per span `key` (ties → REVIEW) in `annotations/adjudicated.jsonl`.
Public app: https://rafmacalaba-data-use-annotate.static.hf.space/