AnesTRACE-Eval

AnesTRACE-Eval is a specialized evaluator for AnesTRACE, a benchmark of clinical reasoning and sequential decision-making in anesthesia and perioperative care. The model assigns structured scores to candidate responses using the AnesTRACE evaluation rubrics.

This repository contains the merged inference checkpoint. It can be loaded directly with Transformers and does not require a separate LoRA adapter.

Intended use

AnesTRACE-Eval is designed for research evaluation of model outputs on AnesTRACE tasks, including:

  • Level Two single-point perioperative decision-making;
  • Level Three multi-turn clinical reasoning and intervention decisions;
  • clinical correctness;
  • evidence-based reasoning and grounding;
  • task completeness;
  • safety severity for intervention-related outputs;
  • temporal adaptation and longitudinal management coherence for multi-turn trajectories.

The model is intended to be used with the official AnesTRACE evaluator prompts and output schemas. It is not intended to generate or validate autonomous clinical care.

Model details

Item Value
Model name AnesTRACE-Eval
Base model Qwen/Qwen3.5-9B
Architecture Qwen3.5 conditional generation model
Primary language English
Parameter precision BF16
Context length in configuration 262,144 tokens
Training framework LLaMA-Factory
Final alignment method Direct Preference Optimization (DPO) with LoRA
Release format Merged safetensors checkpoint

The vision tower was frozen during fine-tuning. The released evaluator is used as a text-based judge in the AnesTRACE evaluation pipeline.

Training

Training was performed in two stages.

Stage 1: supervised evaluator fine-tuning

The base Qwen3.5-9B model was fine-tuned on AnesTRACE evaluator examples using LoRA. The supervised data teach the model the Level Two and Level Three evaluation rubrics, structured score schemas, and safety labels.

Key SFT settings:

Setting Value
LoRA rank 16
LoRA alpha 32
LoRA dropout 0.05
Learning rate 1e-4
Epochs 5
Sequence cutoff 8,192 tokens
Scheduler Cosine
Warmup ratio 0.05
Precision BF16

Stage 2: preference optimization

The merged SFT model was further aligned using true preference pairs derived from Level Three action evaluation data. DPO was applied through a new LoRA adapter, which was subsequently merged into the SFT model.

Key DPO settings:

Setting Value
Objective Sigmoid DPO
Preference beta 0.1
LoRA rank 8
LoRA alpha 16
LoRA dropout 0.05
Learning rate 3e-6
Epochs 1
Sequence cutoff 6,144 tokens
Per-device training batch size 1
Gradient accumulation steps 16
Scheduler Cosine
Warmup ratio 0.05
Seed 42
Precision BF16
Released checkpoint Step 105

The final checkpoint contains the merged model weights in four safetensors shards.

Usage

Install a recent Transformers version with Qwen3.5 support:

pip install -U "transformers>=5.8.0" accelerate safetensors

The following example performs deterministic text-only evaluation. Replace the abbreviated prompts with the official AnesTRACE system prompt and case input for the target evaluation level.

import torch
from transformers import AutoProcessor, Qwen3_5ForConditionalGeneration

model_id = "DataXAI/AnesTRACE-Eval"

processor = AutoProcessor.from_pretrained(
    model_id,
    trust_remote_code=True,
)
model = Qwen3_5ForConditionalGeneration.from_pretrained(
    model_id,
    torch_dtype=torch.bfloat16,
    device_map="auto",
    trust_remote_code=True,
)

messages = [
    {
        "role": "system",
        "content": "You are the AnesTRACE clinical evaluation model. Follow the supplied rubric and return only the required JSON object.",
    },
    {
        "role": "user",
        "content": "[Evaluation Instruction]\nEvaluate the candidate response using the official AnesTRACE rubric.\n\n[Case]\n...\n\n[Reference Answer]\n...\n\n[Candidate Answer]\n...",
    },
]

inputs = processor.apply_chat_template(
    messages,
    tokenize=True,
    add_generation_prompt=True,
    return_dict=True,
    return_tensors="pt",
    enable_thinking=False,
)
inputs = inputs.to(model.device)

with torch.inference_mode():
    generated = model.generate(
        **inputs,
        max_new_tokens=2048,
        do_sample=False,
    )

new_tokens = generated[:, inputs["input_ids"].shape[1]:]
output = processor.batch_decode(
    new_tokens,
    skip_special_tokens=True,
    clean_up_tokenization_spaces=False,
)[0]
print(output)

For reproducible benchmark scoring, use the complete official prompt for the selected level, disable sampling, validate the returned JSON against the corresponding schema, and derive aggregate totals from the validated dimension scores.

Output interpretation

AnesTRACE-Eval uses discrete rubric scores. For Level Two, each B1–B4 task is evaluated independently on:

  • d1_clinical_correctness;
  • d2_evidence_based_reasoning;
  • d3_task_completeness.

Each dimension uses an integer score from 0 to 2. Intervention Decision and Reassessment Plan additionally receive a safety severity label: safe, minor, major, or critical.

Level Three turn evaluation applies the same three dimensions to diagnosis and intervention decisions. A separate trajectory evaluation measures temporal evidence and response adaptation, and longitudinal management coherence.

The evaluator's raw output should be schema-validated before scores are aggregated. Application code should recompute total scores from validated dimension scores rather than trusting a model-generated total.

Limitations

  • AnesTRACE-Eval is a learned evaluator and can make scoring or calibration errors.
  • Scores may be sensitive to prompt formatting, missing reference evidence, truncated candidate answers, and outputs that do not follow the expected schema.
  • The model was optimized for the English AnesTRACE evaluation format. Performance on other languages, unrelated medical specialties, or arbitrary free-form judging tasks has not been established.
  • Agreement with this evaluator does not establish clinical correctness or patient safety.
  • The model must not be used as a medical device, for diagnosis or treatment, or as the sole basis for clinical, regulatory, or deployment decisions.
  • High-stakes results should be reviewed by qualified clinicians, with disagreement analysis and human adjudication where appropriate.

Data and privacy

The model is intended for evaluation on de-identified research data. Users are responsible for ensuring that inputs comply with applicable privacy, institutional, and data-governance requirements. Do not submit identifiable patient information.

License

This model is released under the Apache 2.0 license, subject to the terms and restrictions of the Qwen3.5 base model and any applicable AnesTRACE data licenses.

Citation

If you use AnesTRACE-Eval, please cite our work.

@misc{huang2026anestrace,
  title={AnesTRACE: Benchmarking Intraoperative Anesthesia from Multimodal Perception to Multi-step Decision-Making},
  author={Huang, Ziwei and Gao, Qi and Ji, Zhe and Yao, Yuanyuan and Zhang, Fengjiang and Yan, Min and Xie, Zhongle and Chen, Gang},
  year={2026},
  eprint={2609.32740},
  archivePrefix={arXiv},
  url={https://arxiv.org/abs/2609.32740}
}

Acknowledgements

AnesTRACE-Eval was developed using Qwen3.5 and LLaMA-Factory. We thank the developers and maintainers of these projects.

Downloads last month
174
Safetensors
Model size
9B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for DataXAI/AnesTRACE-Eval

Finetuned
Qwen/Qwen3.5-9B
Finetuned
(993)
this model

Paper for DataXAI/AnesTRACE-Eval