TraceML Labelers (Qwen3-1.7B)

📄 Paper · 🤗 Dataset · 💻 Toolkit · 🌐 Project page

The two labelers of TraceML (NeurIPS 2026, Evaluations & Datasets Track), one per subfolder. Both are Qwen3-1.7B fine-tuned on schema-constrained labels from a larger GPT teacher model, and together they labeled all 151,088 code versions in TraceML.

Subfolder Input Output (JSON)
state/ one version of an ML solution's code the ML-pipeline stages it contains: 8 coarse tags, fine tags from a closed list of 136 with a confidence each, a summary and keywords
action/ a transition between two versions: the code diff, both state labels and the score change what the edit did and why: 10 coarse actions, fine actions from a closed list of 85, one or two of 6 intents, the edit magnitude and the score effect

Use

The simplest route is the TraceML toolkit. It rebuilds the exact prompts from a run, fits long code into the context window, decodes greedily and parses the output:

pip install "traceml-toolkit[label] @ git+https://github.com/JerryYan123/TraceML"
traceml analyze runs/my_run    # extract every code version, label states and transitions, report against human cohorts

Direct use of the state labeler with transformers; the prompt builders and the output parser come from the toolkit:

import torch
from transformers import AutoModelForCausalLM, AutoTokenizer
from traceml_toolkit.labeling import render
from traceml_toolkit.labeling.parse import parse_state_output

repo, sub = "jerryyan/TraceML-Labelers", "state"
device = "cuda" if torch.cuda.is_available() else "mps" if torch.backends.mps.is_available() else "cpu"
tok = AutoTokenizer.from_pretrained(repo, subfolder=sub)
model = AutoModelForCausalLM.from_pretrained(
    repo, subfolder=sub, torch_dtype=torch.float32 if device == "cpu" else torch.bfloat16).to(device)

code = open("train.py").read()
rec = {"comp": "commonlitreadabilityprize", "group": "my-agent", "version_number": 1,
       "code_text": code, "code_lines": code.count("\n") + 1}
prompt = render.render(tok, render.system_prompt("state"), render.user_prompt("state", rec))
inputs = tok(prompt, return_tensors="pt", add_special_tokens=False).to(device)
out = model.generate(**inputs, max_new_tokens=2000, do_sample=False, temperature=None, top_p=None, top_k=None)
print(parse_state_output(tok.decode(out[0, inputs["input_ids"].shape[1]:], skip_special_tokens=True)))

Each transition needs the state labels of both versions, so the action labeler runs after the state labeler; traceml analyze handles the order. To call it yourself, traceml_toolkit.labeling.inputs.build_action_records builds the input records and traceml_toolkit.labeling.prompts.action the prompts and the parser. vLLM takes a local directory, so fetch one labeler first with snapshot_download("jerryyan/TraceML-Labelers", allow_patterns="state/*").

Decode greedily with thinking disabled, as above. generation_config.json keeps Qwen3's sampling defaults, so pass the greedy settings explicitly. The labels released in TraceML were produced with vLLM 0.8.5 in bf16, greedy decoding and prefix caching disabled.

Output

State labeler:

{"coarse_tags": ["data_io", "feature_eng", "model_def", "training_cfg", "validation_cv"],
 "fine_tags": [{"tag": "...", "parent": "model_def", "confidence": "high"}],
 "summary": "...", "keywords": ["..."]}

Action labeler:

{"coarse_actions": ["model", "training"],
 "fine_actions": [{"action": "...", "parent": "model", "confidence": "high"}],
 "intents": [{"intent": "optimization", "confidence": "high"}],
 "goal_nl": "...", "diff_summary": "...", "magnitude": "minor", "score_effect": "improving"}
Field Values
state coarse tags data_io, feature_eng, model_def, training_cfg, ensemble_blend, validation_cv, inference_submit, infra_util
coarse actions data, features, augmentation, model, training, ensemble, validation, inference, infra, housekeeping
intents exploration, optimization, pivoting, debugging, restructuring, verification
magnitude micro, minor, major, overhaul
score effect improving, plateau, regressing, unknown

The full vocabularies are in manifests/schemas/ of the dataset.

Files

state/ and action/ hold the same files as models/qwen3-1.7b-{state,action}/final/ in the dataset repository, without training_args.bin.

License

Apache-2.0, inherited from Qwen3.

Citation

@inproceedings{yan2026traceml,
  title         = {TraceML: What Auto-Research Agents Miss in Long-Horizon ML Development},
  author        = {Yan, Jiarui and Sun, Weiwei and Li, Sijie and Li, Wenhan and Yang, Yiming},
  booktitle     = {Advances in Neural Information Processing Systems (NeurIPS), Track on Evaluations and Datasets},
  year          = {2026},
  eprint        = {2608.26086},
  archivePrefix = {arXiv}
}
Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for jerryyan/TraceML-Labelers

Finetuned
Qwen/Qwen3-1.7B
Finetuned
(1225)
this model

Dataset used to train jerryyan/TraceML-Labelers

Paper for jerryyan/TraceML-Labelers