TraceML Action Labeler (Qwen3-1.7B)

Moved: both labelers now live in one repo, jerryyan/TraceML-Labelers (state/ and action/). This repo is kept only so existing links keep working.

📄 Paper · 🤗 Dataset · 💻 Toolkit · 🌐 Project page · State labeler

The action labeler of TraceML (NeurIPS 2026, Evaluations & Datasets Track). Given a transition between two versions of an ML solution (the code diff, the state labels of both versions and the score change), it returns as JSON what the edit did and why: 10 coarse actions, fine actions from a closed list of 85, one or two intents out of 6, the edit magnitude, the score effect, and short natural-language summaries.

It is Qwen3-1.7B fine-tuned on schema-constrained labels from a larger GPT teacher model. Together with the state labeler it labeled all 151,088 code versions in TraceML.

Use

The simplest route is the TraceML toolkit. It rebuilds the exact prompts from a run, fits long code into the context window, decodes greedily and parses the output:

pip install "traceml-toolkit[label] @ git+https://github.com/JerryYan123/TraceML"
traceml analyze runs/my_run    # extract every code version, label states and transitions, report against human cohorts

Each transition needs the state labels of both versions, so the action labeler runs after the state labeler; traceml analyze handles the order. The toolkit's traceml_toolkit.labeling.inputs.build_action_records builds the input records and traceml_toolkit.labeling.prompts.action the prompts and the parser, if you want to call the model yourself.

Decode greedily with thinking disabled, as above. generation_config.json keeps Qwen3's sampling defaults, so pass the greedy settings explicitly. The labels released in TraceML were produced with vLLM 0.8.5 in bf16, greedy decoding and prefix caching disabled.

Output

{"coarse_actions": ["model", "training"],
 "fine_actions": [{"action": "...", "parent": "model", "confidence": "high"}],
 "intents": [{"intent": "optimization", "confidence": "high"}],
 "goal_nl": "...", "diff_summary": "...", "magnitude": "minor", "score_effect": "improving"}
Field Values
coarse actions data, features, augmentation, model, training, ensemble, validation, inference, infra, housekeeping
intents exploration, optimization, pivoting, debugging, restructuring, verification
magnitude micro, minor, major, overhaul
score effect improving, plateau, regressing, unknown

The full vocabulary is in manifests/schemas/ of the dataset.

Files

The same files as models/qwen3-1.7b-action/final/ in the dataset repository, without training_args.bin.

License

Apache-2.0, inherited from Qwen3.

Citation

@inproceedings{yan2026traceml,
  title         = {TraceML: What Auto-Research Agents Miss in Long-Horizon ML Development},
  author        = {Yan, Jiarui and Sun, Weiwei and Li, Sijie and Li, Wenhan and Yang, Yiming},
  booktitle     = {Advances in Neural Information Processing Systems (NeurIPS), Track on Evaluations and Datasets},
  year          = {2026},
  eprint        = {2608.26086},
  archivePrefix = {arXiv}
}
Downloads last month
-
Safetensors
Model size
2B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for jerryyan/TraceML-Action-Labeler

Finetuned
Qwen/Qwen3-1.7B
Finetuned
(1225)
this model

Dataset used to train jerryyan/TraceML-Action-Labeler

Paper for jerryyan/TraceML-Action-Labeler