TraceML State Labeler (Qwen3-1.7B)

Moved: both labelers now live in one repo, jerryyan/TraceML-Labelers (state/ and action/). This repo is kept only so existing links keep working.

📄 Paper · 🤗 Dataset · 💻 Toolkit · 🌐 Project page · Action labeler

The state labeler of TraceML (NeurIPS 2026, Evaluations & Datasets Track). Given one version of an ML solution's code, it returns the ML-pipeline stages the code contains as JSON: 8 coarse tags, fine tags from a closed list of 136 with a confidence each, a one-line summary and keywords.

It is Qwen3-1.7B fine-tuned on schema-constrained labels from a larger GPT teacher model. Together with the action labeler it labeled all 151,088 code versions in TraceML.

Use

The simplest route is the TraceML toolkit. It rebuilds the exact prompts from a run, fits long code into the context window, decodes greedily and parses the output:

pip install "traceml-toolkit[label] @ git+https://github.com/JerryYan123/TraceML"
traceml analyze runs/my_run    # extract every code version, label states and transitions, report against human cohorts

Direct use with transformers; the prompt builders and the output parser come from the toolkit:

import torch
from transformers import AutoModelForCausalLM, AutoTokenizer
from traceml_toolkit.labeling import render
from traceml_toolkit.labeling.parse import parse_state_output

repo = "jerryyan/TraceML-State-Labeler"
device = "cuda" if torch.cuda.is_available() else "mps" if torch.backends.mps.is_available() else "cpu"
tok = AutoTokenizer.from_pretrained(repo)
model = AutoModelForCausalLM.from_pretrained(
    repo, torch_dtype=torch.float32 if device == "cpu" else torch.bfloat16).to(device)

code = open("train.py").read()
rec = {"comp": "commonlitreadabilityprize", "group": "my-agent", "version_number": 1,
       "code_text": code, "code_lines": code.count("\n") + 1}
prompt = render.render(tok, render.system_prompt("state"), render.user_prompt("state", rec))
inputs = tok(prompt, return_tensors="pt", add_special_tokens=False).to(device)
out = model.generate(**inputs, max_new_tokens=2000, do_sample=False, temperature=None, top_p=None, top_k=None)
print(parse_state_output(tok.decode(out[0, inputs["input_ids"].shape[1]:], skip_special_tokens=True)))

Decode greedily with thinking disabled, as above. generation_config.json keeps Qwen3's sampling defaults, so pass the greedy settings explicitly. The labels released in TraceML were produced with vLLM 0.8.5 in bf16, greedy decoding and prefix caching disabled.

Output

{"coarse_tags": ["data_io", "feature_eng", "model_def", "training_cfg", "validation_cv"],
 "fine_tags": [{"tag": "...", "parent": "model_def", "confidence": "high"}],
 "summary": "...", "keywords": ["..."]}

The coarse tags are data_io, feature_eng, model_def, training_cfg, ensemble_blend, validation_cv, inference_submit and infra_util. The full vocabulary is in manifests/schemas/ of the dataset.

Files

The same files as models/qwen3-1.7b-state/final/ in the dataset repository, without training_args.bin.

License

Apache-2.0, inherited from Qwen3.

Citation

@inproceedings{yan2026traceml,
  title         = {TraceML: What Auto-Research Agents Miss in Long-Horizon ML Development},
  author        = {Yan, Jiarui and Sun, Weiwei and Li, Sijie and Li, Wenhan and Yang, Yiming},
  booktitle     = {Advances in Neural Information Processing Systems (NeurIPS), Track on Evaluations and Datasets},
  year          = {2026},
  eprint        = {2608.26086},
  archivePrefix = {arXiv}
}
Downloads last month
-
Safetensors
Model size
2B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for jerryyan/TraceML-State-Labeler

Finetuned
Qwen/Qwen3-1.7B
Finetuned
(1225)
this model

Dataset used to train jerryyan/TraceML-State-Labeler

Paper for jerryyan/TraceML-State-Labeler