RLHarness MapTab Models

This repository contains the two fully merged Qwen3.5-9B checkpoints released with RLHarness for multimodal route planning on the MetroMap and TravelMap domains of MapTab. The checkpoints combine the Qwen3.5-9B model, domain-specific supervised fine-tuning, and the selected second-round reinforcement-learning update. They are ready for inference and do not require a separate LoRA adapter.

Model variants

Hub subfolder Domain Format Reported Test400 exact accuracy
metromap_trained/ MetroMap BF16 safetensors, 4 shards 62.00%
travelmap_trained/ TravelMap BF16 safetensors, 4 shards 50.00%

Each checkpoint contains approximately 18.8 GB of model weights. The TravelMap run additionally records 58.48% partial route accuracy, a 98.00% valid-format rate, and 8 truncated responses among 400 examples in travelmap_trained/travelmap_test400_summary.json.

The two subfolders are independent checkpoints. Select the checkpoint matching the input domain; do not combine them and do not attach the original training adapter again.

Download and load one checkpoint

Install a recent Transformers release with Qwen3.5 support:

python -m pip install "transformers>=5.10.0" "huggingface_hub>=0.32" accelerate safetensors

Download only one domain instead of the full repository:

from huggingface_hub import snapshot_download

repo_id = "szq-nju/RLHarness-MapTab-Models"
subfolder = "metromap_trained"  # or "travelmap_trained"

snapshot_download(
    repo_id=repo_id,
    local_dir="RLHarness-MapTab-Models",
    allow_patterns=[f"{subfolder}/*"],
)
model_path = f"RLHarness-MapTab-Models/{subfolder}"

Load the downloaded model with Transformers:

from transformers import AutoModelForImageTextToText, AutoProcessor

processor = AutoProcessor.from_pretrained(model_path, trust_remote_code=True)
model = AutoModelForImageTextToText.from_pretrained(
    model_path,
    dtype="auto",
    device_map="auto",
    trust_remote_code=True,
)

The checkpoint can also be loaded directly from the Hub with its subfolder:

from transformers import AutoModelForImageTextToText, AutoProcessor

repo_id = "szq-nju/RLHarness-MapTab-Models"
subfolder = "travelmap_trained"

processor = AutoProcessor.from_pretrained(repo_id, subfolder=subfolder)
model = AutoModelForImageTextToText.from_pretrained(
    repo_id,
    subfolder=subfolder,
    dtype="auto",
    device_map="auto",
)

End-to-end RLHarness evaluation

The RLHarness framework can download the selected subfolder, verify the model index and all referenced shards, apply the released domain prompt, and evaluate the locked Test400 split with vLLM:

git clone https://github.com/Ziqiao-Shang/RLHarness.git
cd RLHarness
python -m pip install -e .

VLLM_PYTHON=/path/to/vllm-env/bin/python \
GPU_IDS="0 1 2 3" \
bash scripts/07_evaluate_pretrained.sh --domain metromap

VLLM_PYTHON=/path/to/vllm-env/bin/python \
GPU_IDS="0 1 2 3" \
bash scripts/07_evaluate_pretrained.sh --domain travelmap

By default, models are cached under models/release/. Set RLHARNESS_MODEL_ROOT to use another location. Set RLHARNESS_MODEL_REVISION to a Hub commit hash for immutable reproduction. RLHarness can obtain the fixed Test400 inputs from the complete official MapTab release by locked sample ID:

bash scripts/prepare_data.sh --download --domain travelmap

For a quick interface check, bash scripts/prepare_data.sh --smoke-test --domain travelmap downloads only one locked train example and one locked test example plus their image/table assets. An already downloaded merged checkpoint can be validated without contacting the model Hub or starting vLLM:

MODEL_PATH=/path/to/travelmap_trained \
bash scripts/07_evaluate_pretrained.sh --domain travelmap --check-only

Training procedure

Both variants use the same staged RLHarness workflow with domain-specific prompts, data, and selected updates:

  1. Build an initial skill-augmented Harness and verified teacher trajectories.
  2. Fine-tune Qwen3.5-9B for two epochs with a rank-16 LoRA adapter.
  3. Run the first GRPO/Hybrid-DGPO stage through step 400 with eight rollouts per question.
  4. Reconstruct and select the Harness using fresh success/failure rollouts and the fixed Val100 set.
  5. Run the second reinforcement-learning stage through step 600.
  6. Merge the selected language LoRA update into the SFT model and evaluate once on the held-out Test400 set.

The fixed split contains 1,600 training, 100 validation, and 400 test examples per domain. Test400 is not used for training, Harness reconstruction, or candidate selection. TravelMap additionally isolates the splits by map.

Intended use

These checkpoints are intended for research on multimodal route planning, skill-augmented reasoning, RL-trained vision-language models, and reproduction of the RLHarness MapTab experiments. Inputs are expected to follow the released MetroMap or TravelMap task contracts, including the map image, vertex table, constraints, and domain prompt.

They are not general navigation systems. The models can produce invalid, incomplete, or hallucinated routes, especially for maps and schemas outside the training distribution. Do not use their outputs for safety-critical transportation or real-world routing without independent validation.

Evaluation notes

  • Primary metric: exact route accuracy on the fixed 400-example test split.
  • Partial accuracy measures overlap with the reference route and is diagnostic rather than the primary selection metric.
  • Native thinking was disabled and generation was capped at 4,096 new tokens for the reported evaluation.
  • Scores depend on the released final prompt and RLHarness preprocessing/evaluation code; generic prompting may produce different results.

Files and provenance

Each model directory contains its model configuration, processor/tokenizer files, chat template, generation configuration, safetensors index, and four safetensors shards. merge_manifest.json records merge provenance. metromap_trained/SHA256SUMS provides checksums for the MetroMap release files.

The base model is Qwen/Qwen3.5-9B, released under Apache 2.0. MapTab data and benchmark details are available from the MapTab repository and MapTab dataset. Users are responsible for following the terms that apply to the base model and data.

Citation

Please cite MapTab when using these checkpoints for the benchmark:

@article{shang2026maptab,
  title   = {MapTab: A Diagnostic Benchmark for Long-Horizon Multi-Criteria Multimodal Reasoning on Heterogeneous Topological Graphs},
  author  = {Shang, Ziqiao and Ge, Lingyue and Xu, Zian and Cheng, Zi-Jian and Tian, Shi-Yu and Huang, Zhenyu and Fu, Wenbo and Wu, Weiming and Chen, Yang and Zhang, Xiangwen and Hu, Yulan and Liu, Bin and Guo, Lan-Zhe},
  journal = {arXiv preprint arXiv:2602.18600},
  year    = {2026}
}
Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for szq-nju/RLHarness-MapTab-Models

Finetuned
Qwen/Qwen3.5-9B
Finetuned
(983)
this model

Paper for szq-nju/RLHarness-MapTab-Models