Instructions to use szq-nju/RLHarness-MapTab-Models with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use szq-nju/RLHarness-MapTab-Models with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("image-text-to-text", model="szq-nju/RLHarness-MapTab-Models")# Load model directly from transformers import AutoModel model = AutoModel.from_pretrained("szq-nju/RLHarness-MapTab-Models", device_map="auto") - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use szq-nju/RLHarness-MapTab-Models with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "szq-nju/RLHarness-MapTab-Models" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "szq-nju/RLHarness-MapTab-Models", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker
docker model run hf.co/szq-nju/RLHarness-MapTab-Models
- SGLang
How to use szq-nju/RLHarness-MapTab-Models with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "szq-nju/RLHarness-MapTab-Models" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "szq-nju/RLHarness-MapTab-Models", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "szq-nju/RLHarness-MapTab-Models" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "szq-nju/RLHarness-MapTab-Models", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }' - Docker Model Runner
How to use szq-nju/RLHarness-MapTab-Models with Docker Model Runner:
docker model run hf.co/szq-nju/RLHarness-MapTab-Models
RLHarness MapTab Models
This repository contains the two fully merged Qwen3.5-9B checkpoints released with RLHarness for multimodal route planning on the MetroMap and TravelMap domains of MapTab. The checkpoints combine the Qwen3.5-9B model, domain-specific supervised fine-tuning, and the selected second-round reinforcement-learning update. They are ready for inference and do not require a separate LoRA adapter.
Model variants
| Hub subfolder | Domain | Format | Reported Test400 exact accuracy |
|---|---|---|---|
metromap_trained/ |
MetroMap | BF16 safetensors, 4 shards | 62.00% |
travelmap_trained/ |
TravelMap | BF16 safetensors, 4 shards | 50.00% |
Each checkpoint contains approximately 18.8 GB of model weights. The TravelMap run additionally records 58.48% partial route accuracy, a 98.00% valid-format rate, and 8 truncated responses among 400 examples in travelmap_trained/travelmap_test400_summary.json.
The two subfolders are independent checkpoints. Select the checkpoint matching the input domain; do not combine them and do not attach the original training adapter again.
Download and load one checkpoint
Install a recent Transformers release with Qwen3.5 support:
python -m pip install "transformers>=5.10.0" "huggingface_hub>=0.32" accelerate safetensors
Download only one domain instead of the full repository:
from huggingface_hub import snapshot_download
repo_id = "szq-nju/RLHarness-MapTab-Models"
subfolder = "metromap_trained" # or "travelmap_trained"
snapshot_download(
repo_id=repo_id,
local_dir="RLHarness-MapTab-Models",
allow_patterns=[f"{subfolder}/*"],
)
model_path = f"RLHarness-MapTab-Models/{subfolder}"
Load the downloaded model with Transformers:
from transformers import AutoModelForImageTextToText, AutoProcessor
processor = AutoProcessor.from_pretrained(model_path, trust_remote_code=True)
model = AutoModelForImageTextToText.from_pretrained(
model_path,
dtype="auto",
device_map="auto",
trust_remote_code=True,
)
The checkpoint can also be loaded directly from the Hub with its subfolder:
from transformers import AutoModelForImageTextToText, AutoProcessor
repo_id = "szq-nju/RLHarness-MapTab-Models"
subfolder = "travelmap_trained"
processor = AutoProcessor.from_pretrained(repo_id, subfolder=subfolder)
model = AutoModelForImageTextToText.from_pretrained(
repo_id,
subfolder=subfolder,
dtype="auto",
device_map="auto",
)
End-to-end RLHarness evaluation
The RLHarness framework can download the selected subfolder, verify the model index and all referenced shards, apply the released domain prompt, and evaluate the locked Test400 split with vLLM:
git clone https://github.com/Ziqiao-Shang/RLHarness.git
cd RLHarness
python -m pip install -e .
VLLM_PYTHON=/path/to/vllm-env/bin/python \
GPU_IDS="0 1 2 3" \
bash scripts/07_evaluate_pretrained.sh --domain metromap
VLLM_PYTHON=/path/to/vllm-env/bin/python \
GPU_IDS="0 1 2 3" \
bash scripts/07_evaluate_pretrained.sh --domain travelmap
By default, models are cached under models/release/. Set RLHARNESS_MODEL_ROOT to use another location. Set RLHARNESS_MODEL_REVISION to a Hub commit hash for immutable reproduction. RLHarness can obtain the fixed Test400 inputs from the complete official MapTab release by locked sample ID:
bash scripts/prepare_data.sh --download --domain travelmap
For a quick interface check, bash scripts/prepare_data.sh --smoke-test --domain travelmap downloads only one locked train example and one locked test example plus their image/table assets. An already downloaded merged checkpoint can be validated without contacting the model Hub or starting vLLM:
MODEL_PATH=/path/to/travelmap_trained \
bash scripts/07_evaluate_pretrained.sh --domain travelmap --check-only
Training procedure
Both variants use the same staged RLHarness workflow with domain-specific prompts, data, and selected updates:
- Build an initial skill-augmented Harness and verified teacher trajectories.
- Fine-tune Qwen3.5-9B for two epochs with a rank-16 LoRA adapter.
- Run the first GRPO/Hybrid-DGPO stage through step 400 with eight rollouts per question.
- Reconstruct and select the Harness using fresh success/failure rollouts and the fixed Val100 set.
- Run the second reinforcement-learning stage through step 600.
- Merge the selected language LoRA update into the SFT model and evaluate once on the held-out Test400 set.
The fixed split contains 1,600 training, 100 validation, and 400 test examples per domain. Test400 is not used for training, Harness reconstruction, or candidate selection. TravelMap additionally isolates the splits by map.
Intended use
These checkpoints are intended for research on multimodal route planning, skill-augmented reasoning, RL-trained vision-language models, and reproduction of the RLHarness MapTab experiments. Inputs are expected to follow the released MetroMap or TravelMap task contracts, including the map image, vertex table, constraints, and domain prompt.
They are not general navigation systems. The models can produce invalid, incomplete, or hallucinated routes, especially for maps and schemas outside the training distribution. Do not use their outputs for safety-critical transportation or real-world routing without independent validation.
Evaluation notes
- Primary metric: exact route accuracy on the fixed 400-example test split.
- Partial accuracy measures overlap with the reference route and is diagnostic rather than the primary selection metric.
- Native thinking was disabled and generation was capped at 4,096 new tokens for the reported evaluation.
- Scores depend on the released final prompt and RLHarness preprocessing/evaluation code; generic prompting may produce different results.
Files and provenance
Each model directory contains its model configuration, processor/tokenizer files, chat template, generation configuration, safetensors index, and four safetensors shards. merge_manifest.json records merge provenance. metromap_trained/SHA256SUMS provides checksums for the MetroMap release files.
The base model is Qwen/Qwen3.5-9B, released under Apache 2.0. MapTab data and benchmark details are available from the MapTab repository and MapTab dataset. Users are responsible for following the terms that apply to the base model and data.
Citation
Please cite MapTab when using these checkpoints for the benchmark:
@article{shang2026maptab,
title = {MapTab: A Diagnostic Benchmark for Long-Horizon Multi-Criteria Multimodal Reasoning on Heterogeneous Topological Graphs},
author = {Shang, Ziqiao and Ge, Lingyue and Xu, Zian and Cheng, Zi-Jian and Tian, Shi-Yu and Huang, Zhenyu and Fu, Wenbo and Wu, Weiming and Chen, Yang and Zhang, Xiangwen and Hu, Yulan and Liu, Bin and Guo, Lan-Zhe},
journal = {arXiv preprint arXiv:2602.18600},
year = {2026}
}