lfm2.5-350M-focus-metadata

LFM2.5 SFT for passage-grounded source attributes and data-use metadata. Given a passage chunk and one cited data mention, the model answers one of two tasks as JSON.

Tasks

Task Target keys Allowed values
attributes tags, attributes tags: geospatial, forced_displacement, agriculture, employment, poverty, education, health, climate. attributes: geography, year, version, producer, acronym (verbatim passage text or null), data_type (survey, census, administrative, microdata, database, indicator, geospatial, report, other)
usage role, usage_action, summary role: primary, supporting, background. usage_action: source, analyze, curate, validate, inform. summary: one sentence in English, whatever language the passage is in

The mention is supplied in the prompt; the model labels that mention, it does not detect mentions. Summaries are English by convention - every annotated row is English, so the prompt asks for English rather than the passage's language.

Training

  • Base model: LiquidAI/LFM2.5-350M
  • Dataset: rafmacalaba/data-use-focus-metadata-sft (attributes + usage configs)
  • Epochs: 5 | LoRA r=16, alpha=32, dropout=0.05 | loss on assistant completion only
  • All published train rows are used; the builder upsamples rare usage_action classes in train to counter lane label drift (see dataset README).
  • Splits are document-disjoint and carry over from the source annotation lanes.
  • Annotation is general focus, agent-generated with verbatim checks; not human gold.

Data used (rows; one mention span yields one row per task)

Split attributes usage Total
train 3832 6522 10354
val 1780 1780 3560
holdout 1588 1588 3176

Holdout performance

Every holdout row was generated and scored; nothing is capped or subsampled. Metrics compare generations against the annotated holdout labels, so they measure agreement with agent annotations, not independent human-gold accuracy.

Evaluated rows: 3176 (attributes 1588, usage 1588)

Attributes task

Metric Value
Valid JSON rate 0.9994
Exact target accuracy 0.0901
Theme tags micro-F1 0.6383
Theme tags exact accuracy 0.4018
data_type macro-F1 0.4319

Per field. accuracy counts correct nulls, so a field that is almost always absent looks solved while recall on the values that exist says otherwise - read recall and support.

Field Support Recall Accuracy (incl. nulls) Null accuracy
geography 827 0.6167 0.6977 0.7858
year 724 0.7541 0.7758 0.7940
version 33 0.0000 0.9792 1.0000
producer 610 0.5131 0.6851 0.7924
acronym 679 0.7938 0.8596 0.9087
data_type 1588 0.7128 0.7128 n/a

Usage task

Metric Value
role + usage_action pair accuracy 0.2809
role accuracy 0.6889
role macro-F1 0.3458
usage_action accuracy 0.3848
usage_action macro-F1 0.2213
summary token-F1 (bag-of-words overlap) 0.2481

summary is free text, so exact string match is not a quality metric and is not reported here; it is kept in holdout_metrics.json only as a strict-copy check. Token-F1 measures lexical overlap, which under-counts valid paraphrase - read a sample of holdout_predictions.jsonl before judging summary quality.

Inference

ChatML, two tasks. The user turn is identical for both: Text: <passage>\n\nMention: "<span>". Only the system prompt selects the task, so send the exact system prompt below.

import json
from transformers import AutoModelForCausalLM, AutoTokenizer

MODEL = "rafmacalaba/lfm2.5-350M-focus-metadata"   # merged weights are at the repo root
tok = AutoTokenizer.from_pretrained(MODEL)
model = AutoModelForCausalLM.from_pretrained(MODEL, trust_remote_code=True,
                                          dtype="bfloat16", device_map="auto")

# Adapter-only variant: load the base model, then the adapter in `adapter/`.
#   from peft import PeftModel
#   base = AutoModelForCausalLM.from_pretrained("LiquidAI/LFM2.5-350M",
#                                              trust_remote_code=True, dtype="bfloat16")
#   model = PeftModel.from_pretrained(base, MODEL, subfolder="adapter")

ATTRIBUTES_SYSTEM = """Extract theme tags and source attributes for the cited data mention using only the passage. Output ONLY JSON with keys tags and attributes. tags is an array from geospatial, forced_displacement, agriculture, employment, poverty, education, health, climate. attributes has geography, year, version, producer, acronym (exact passage text or null) and data_type (survey, census, administrative, microdata, database, indicator, geospatial, report, or other). Do not infer values absent from the passage."""
USAGE_SYSTEM = """Describe how the author uses the cited data mention, using only the passage. Output ONLY JSON with keys role, usage_action, summary. Pick usage_action from these boundaries: source = the mention is where the quoted figures come from; analyze = the author computes, estimates, regresses or aggregates values from it; validate = it is used to cross-check or corroborate another estimate; curate = the author describes collecting, compiling, archiving or maintaining it; inform = cited for context or trend without its values being used. role is primary, supporting, or background. summary is one concise sentence in English, whatever language the passage is in; do not invent details."""

def predict(system, passage, span, max_new_tokens=256):
    messages = [
        {"role": "system", "content": system},
        {"role": "user",
         "content": f'Text: {passage}\n\nMention: "{span}"'},
    ]
    ids = tok.apply_chat_template(messages, add_generation_prompt=True,
                                  return_dict=False, return_tensors="pt")[0]
    ids = ids.unsqueeze(0).to(model.device)
    out = model.generate(ids, max_new_tokens=max_new_tokens, do_sample=False)
    raw = tok.decode(out[0][ids.shape[1]:], skip_special_tokens=True)
    # LFM2.5 is a thinking model: it may emit <|im_start|> reasoning before the JSON.
    answer = raw.split("<|im_start|>")[-1] if "<|im_start|>" in raw else raw
    start, end = answer.find("{"), answer.rfind("}")
    return json.loads(answer[start:end + 1])

passage = ("Figure A1: Monthly Income by Education, 2018 Household Survey")
span = "2018 Household Survey"

print(predict(ATTRIBUTES_SYSTEM, passage, span))
# {'tags': ['employment', 'education'], 'attributes': {'geography': None, 'year': '2018',
#  'version': None, 'producer': None, 'acronym': None, 'data_type': 'survey'}}

print(predict(USAGE_SYSTEM, passage, span))
# {'role': 'primary', 'usage_action': 'source',
#  'summary': 'Monthly income by education level is tabulated from this survey.'}

Batch note: pad left and slice the prompt to --max-length (2048) tokens, as in training/finetune_lfm2_focus_metadata.py::evaluate_holdout.

Files

  • holdout_metrics.json - every metric above, machine-readable
  • holdout_predictions.jsonl - per-row task, source_key, occurrence_id, gold, pred
  • trainer_state.json - per-epoch train/eval loss curve from the HF Trainer
  • config.json + weights at the repo root - merged model, loads directly
  • adapter/ - LoRA adapter (needs the base model above; a PEFT config at the repo root would otherwise make subfolder= loads resolve to the base model instead)
Downloads last month
302
Safetensors
Model size
0.4B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support