lfm2.5-350M-focus-metadata
LFM2.5 SFT for passage-grounded source attributes and data-use metadata. Given a passage chunk and one cited data mention, the model answers one of two tasks as JSON.
Tasks
| Task | Target keys | Allowed values |
|---|---|---|
attributes |
tags, attributes |
tags: geospatial, forced_displacement, agriculture, employment, poverty, education, health, climate. attributes: geography, year, version, producer, acronym (verbatim passage text or null), data_type (survey, census, administrative, microdata, database, indicator, geospatial, report, other) |
usage |
role, usage_action, summary |
role: primary, supporting, background. usage_action: source, analyze, curate, validate, inform. summary: one sentence in English, whatever language the passage is in |
The mention is supplied in the prompt; the model labels that mention, it does not detect mentions. Summaries are English by convention - every annotated row is English, so the prompt asks for English rather than the passage's language.
Training
- Base model:
LiquidAI/LFM2.5-350M - Dataset:
rafmacalaba/data-use-focus-metadata-sft(attributes+usageconfigs) - Epochs: 5 | LoRA r=16, alpha=32, dropout=0.05 | loss on assistant completion only
- All published train rows are used; the builder upsamples rare
usage_actionclasses in train to counter lane label drift (see dataset README). - Splits are document-disjoint and carry over from the source annotation lanes.
- Annotation is
generalfocus, agent-generated with verbatim checks; not human gold.
Data used (rows; one mention span yields one row per task)
| Split | attributes |
usage |
Total |
|---|---|---|---|
| train | 3832 | 6522 | 10354 |
| val | 1780 | 1780 | 3560 |
| holdout | 1588 | 1588 | 3176 |
Holdout performance
Every holdout row was generated and scored; nothing is capped or subsampled. Metrics compare generations against the annotated holdout labels, so they measure agreement with agent annotations, not independent human-gold accuracy.
Evaluated rows: 3176 (attributes 1588, usage 1588)
Attributes task
| Metric | Value |
|---|---|
| Valid JSON rate | 0.9994 |
| Exact target accuracy | 0.0901 |
| Theme tags micro-F1 | 0.6383 |
| Theme tags exact accuracy | 0.4018 |
| data_type macro-F1 | 0.4319 |
Per field. accuracy counts correct nulls, so a field that is almost always absent looks solved while recall on the values that exist says otherwise - read recall and support.
| Field | Support | Recall | Accuracy (incl. nulls) | Null accuracy |
|---|---|---|---|---|
geography |
827 | 0.6167 | 0.6977 | 0.7858 |
year |
724 | 0.7541 | 0.7758 | 0.7940 |
version |
33 | 0.0000 | 0.9792 | 1.0000 |
producer |
610 | 0.5131 | 0.6851 | 0.7924 |
acronym |
679 | 0.7938 | 0.8596 | 0.9087 |
data_type |
1588 | 0.7128 | 0.7128 | n/a |
Usage task
| Metric | Value |
|---|---|
role + usage_action pair accuracy |
0.2809 |
role accuracy |
0.6889 |
role macro-F1 |
0.3458 |
usage_action accuracy |
0.3848 |
usage_action macro-F1 |
0.2213 |
summary token-F1 (bag-of-words overlap) |
0.2481 |
summary is free text, so exact string match is not a quality metric and is not reported here; it is kept in holdout_metrics.json only as a strict-copy check. Token-F1 measures lexical overlap, which under-counts valid paraphrase - read a sample of holdout_predictions.jsonl before judging summary quality.
Inference
ChatML, two tasks. The user turn is identical for both: Text: <passage>\n\nMention: "<span>". Only the system prompt selects the task, so send the exact system prompt below.
import json
from transformers import AutoModelForCausalLM, AutoTokenizer
MODEL = "rafmacalaba/lfm2.5-350M-focus-metadata" # merged weights are at the repo root
tok = AutoTokenizer.from_pretrained(MODEL)
model = AutoModelForCausalLM.from_pretrained(MODEL, trust_remote_code=True,
dtype="bfloat16", device_map="auto")
# Adapter-only variant: load the base model, then the adapter in `adapter/`.
# from peft import PeftModel
# base = AutoModelForCausalLM.from_pretrained("LiquidAI/LFM2.5-350M",
# trust_remote_code=True, dtype="bfloat16")
# model = PeftModel.from_pretrained(base, MODEL, subfolder="adapter")
ATTRIBUTES_SYSTEM = """Extract theme tags and source attributes for the cited data mention using only the passage. Output ONLY JSON with keys tags and attributes. tags is an array from geospatial, forced_displacement, agriculture, employment, poverty, education, health, climate. attributes has geography, year, version, producer, acronym (exact passage text or null) and data_type (survey, census, administrative, microdata, database, indicator, geospatial, report, or other). Do not infer values absent from the passage."""
USAGE_SYSTEM = """Describe how the author uses the cited data mention, using only the passage. Output ONLY JSON with keys role, usage_action, summary. Pick usage_action from these boundaries: source = the mention is where the quoted figures come from; analyze = the author computes, estimates, regresses or aggregates values from it; validate = it is used to cross-check or corroborate another estimate; curate = the author describes collecting, compiling, archiving or maintaining it; inform = cited for context or trend without its values being used. role is primary, supporting, or background. summary is one concise sentence in English, whatever language the passage is in; do not invent details."""
def predict(system, passage, span, max_new_tokens=256):
messages = [
{"role": "system", "content": system},
{"role": "user",
"content": f'Text: {passage}\n\nMention: "{span}"'},
]
ids = tok.apply_chat_template(messages, add_generation_prompt=True,
return_dict=False, return_tensors="pt")[0]
ids = ids.unsqueeze(0).to(model.device)
out = model.generate(ids, max_new_tokens=max_new_tokens, do_sample=False)
raw = tok.decode(out[0][ids.shape[1]:], skip_special_tokens=True)
# LFM2.5 is a thinking model: it may emit <|im_start|> reasoning before the JSON.
answer = raw.split("<|im_start|>")[-1] if "<|im_start|>" in raw else raw
start, end = answer.find("{"), answer.rfind("}")
return json.loads(answer[start:end + 1])
passage = ("Figure A1: Monthly Income by Education, 2018 Household Survey")
span = "2018 Household Survey"
print(predict(ATTRIBUTES_SYSTEM, passage, span))
# {'tags': ['employment', 'education'], 'attributes': {'geography': None, 'year': '2018',
# 'version': None, 'producer': None, 'acronym': None, 'data_type': 'survey'}}
print(predict(USAGE_SYSTEM, passage, span))
# {'role': 'primary', 'usage_action': 'source',
# 'summary': 'Monthly income by education level is tabulated from this survey.'}
Batch note: pad left and slice the prompt to --max-length (2048) tokens, as in training/finetune_lfm2_focus_metadata.py::evaluate_holdout.
Files
holdout_metrics.json- every metric above, machine-readableholdout_predictions.jsonl- per-rowtask,source_key,occurrence_id,gold,predtrainer_state.json- per-epoch train/eval loss curve from the HF Trainerconfig.json+ weights at the repo root - merged model, loads directlyadapter/- LoRA adapter (needs the base model above; a PEFT config at the repo root would otherwise makesubfolder=loads resolve to the base model instead)
- Downloads last month
- 302