Instructions to use internlm/AdvancedMathBench-AutoVerifier with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use internlm/AdvancedMathBench-AutoVerifier with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("image-text-to-text", model="internlm/AdvancedMathBench-AutoVerifier") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] pipe(text=messages)# Load model directly from transformers import AutoProcessor, AutoModelForMultimodalLM processor = AutoProcessor.from_pretrained("internlm/AdvancedMathBench-AutoVerifier") model = AutoModelForMultimodalLM.from_pretrained("internlm/AdvancedMathBench-AutoVerifier", device_map="auto") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] inputs = processor.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=40) print(processor.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use internlm/AdvancedMathBench-AutoVerifier with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "internlm/AdvancedMathBench-AutoVerifier" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "internlm/AdvancedMathBench-AutoVerifier", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker
docker model run hf.co/internlm/AdvancedMathBench-AutoVerifier
- SGLang
How to use internlm/AdvancedMathBench-AutoVerifier with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "internlm/AdvancedMathBench-AutoVerifier" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "internlm/AdvancedMathBench-AutoVerifier", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "internlm/AdvancedMathBench-AutoVerifier" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "internlm/AdvancedMathBench-AutoVerifier", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }' - Docker Model Runner
How to use internlm/AdvancedMathBench-AutoVerifier with Docker Model Runner:
docker model run hf.co/internlm/AdvancedMathBench-AutoVerifier
# Load model directly
from transformers import AutoProcessor, AutoModelForMultimodalLM
processor = AutoProcessor.from_pretrained("internlm/AdvancedMathBench-AutoVerifier")
model = AutoModelForMultimodalLM.from_pretrained("internlm/AdvancedMathBench-AutoVerifier", device_map="auto")
messages = [
{
"role": "user",
"content": [
{"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"},
{"type": "text", "text": "What animal is on the candy?"}
]
},
]
inputs = processor.apply_chat_template(
messages,
add_generation_prompt=True,
tokenize=True,
return_dict=True,
return_tensors="pt",
).to(model.device)
outputs = model.generate(**inputs, max_new_tokens=40)
print(processor.decode(outputs[0][inputs["input_ids"].shape[-1]:]))AdvancedMathBench AutoVerifier
Paper · HF Paper · GitHub · Dataset · AutoVerifier
AutoVerifier evaluates natural-language mathematical proofs, explains errors, and identifies the earliest incorrect step. It serves as the automatic grader for AdvancedMathBench's ProverBench.
Model
- Architecture:
Qwen3_5MoeForConditionalGeneration. - Tokenizer: bundled
InternS1Tokenizer; requiressentencepieceandtrust_remote_code=Trueafter reviewing the tokenizer code. - Weights: 40 safetensors shards, approximately 68 GiB.
Input
Use proof_verifier.md with a problem, an optional reference solution, and a candidate proof split into zero-indexed steps. The following constructs the input without loading the model weights:
from pathlib import Path
from transformers import AutoTokenizer
model_dir = "." # Local model repository directory
tokenizer = AutoTokenizer.from_pretrained(model_dir, trust_remote_code=True)
steps = ["A candidate proof step.", "Another candidate proof step."]
proof = "\n\n".join(
f"<step{i}>\n\n{step}\n\n</step{i}>" for i, step in enumerate(steps)
)
template = Path(model_dir, "prompts/proof_verifier.md").read_text(encoding="utf-8")
prompt = template.format(
problem="The mathematical problem.", human_solution="", solution=proof,
)
text = tokenizer.apply_chat_template(
[{"role": "user", "content": prompt}],
tokenize=False, add_generation_prompt=True, enable_thinking=True,
)
Output and scoring
The final response contains an assessment, identified errors, and the first error index. For example, a no-error judgment is:
<assessment>The proof is correct.</assessment>
<errors></errors>
<first_error_step>-1</first_error_step>
-1 means no error was found; nonnegative indices identify the earliest error,
starting from 0. Parse the final answer after </think> when present.
ProverBench checks each proof 8 times and accepts it only when all eight
valid judgments report -1. Missing or malformed judgments do not count as
acceptance. Sampling settings are documented in
evaluation_settings.json.
Notes
AutoVerifier is a learned grader, not a formal proof checker, and can make errors. Tested package versions and validation scope are recorded in compatibility.json. License notices are provided in LICENSE and NOTICE.md.
Citation
If you use AdvancedMathBench AutoVerifier in your research, please cite:
@misc{kong2026advancedmathbenchbenchmarksuiteadvanced,
title = {AdvancedMathBench: A Benchmark Suite for Advanced Mathematical Proof Generation and Verification},
author = {Lingkai Kong and Zijian Wu and Yuzhe Gu and Haiteng Zhao and Wenyong Huang and Shuang Sun and Zhicheng Xiong and Xiaotian Zhang and Shuya Zhao and Yan Wang and Disheng Xu and Wenwei Zhang and Kai Chen},
year = {2026},
eprint = {2607.11849},
archivePrefix = {arXiv},
primaryClass = {cs.CL},
doi = {10.48550/arXiv.2607.11849},
url = {https://arxiv.org/abs/2607.11849}
}
- Downloads last month
- -
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("image-text-to-text", model="internlm/AdvancedMathBench-AutoVerifier") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] pipe(text=messages)