Instructions to use Qwen/Qwen2.5-Math-1.5B-Instruct with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use Qwen/Qwen2.5-Math-1.5B-Instruct with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="Qwen/Qwen2.5-Math-1.5B-Instruct") messages = [ {"role": "user", "content": "Who are you?"}, ] pipe(messages)# pip install -U transformers accelerate # Load model directly from transformers import AutoTokenizer, AutoModelForCausalLM tokenizer = AutoTokenizer.from_pretrained("Qwen/Qwen2.5-Math-1.5B-Instruct") model = AutoModelForCausalLM.from_pretrained("Qwen/Qwen2.5-Math-1.5B-Instruct", device_map="auto") messages = [ {"role": "user", "content": "Who are you?"}, ] inputs = tokenizer.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=256) print(tokenizer.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Inference
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use Qwen/Qwen2.5-Math-1.5B-Instruct with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "Qwen/Qwen2.5-Math-1.5B-Instruct" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Qwen/Qwen2.5-Math-1.5B-Instruct", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/Qwen/Qwen2.5-Math-1.5B-Instruct
- SGLang
How to use Qwen/Qwen2.5-Math-1.5B-Instruct with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "Qwen/Qwen2.5-Math-1.5B-Instruct" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Qwen/Qwen2.5-Math-1.5B-Instruct", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "Qwen/Qwen2.5-Math-1.5B-Instruct" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Qwen/Qwen2.5-Math-1.5B-Instruct", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Docker Model Runner
How to use Qwen/Qwen2.5-Math-1.5B-Instruct with Docker Model Runner:
docker model run hf.co/Qwen/Qwen2.5-Math-1.5B-Instruct
Measurement note: Qwen2.5-1.5B β Qwen2.5-Math-1.5B-Instruct at exact revisions (Model X-Ray, 22 September 2026)
We compared Qwen/Qwen2.5-1.5B @ 8faed761d45a263340a0528343f099c05c9a4323 (A) with this repository @ aafeb0fc6f22cbf0eaeed126eff8be45b0360a35 (B) with our instrument (Tetracta Model X-Ray: instrument VG1; measurement contract mv-1.4; report schema rs-1.7 / presentation rp-1.3) on 22 September 2026. The report is labelled a validation-pending descriptive result. We are posting it here so that the measurement is on record next to the artifact it describes and can be corrected by people who know these checkpoints better than we do.
What the report states, and nothing more:
- Internal response: a difference was observed at 25 of 25 evaluated positions (decoder block 4 output through the final normalized output). The embedding output and blocks 1β3 are not evaluated, so where the difference begins is unresolved; "not evaluated" is not "no difference".
- Text output: withheld β the two artifacts do not share a directly comparable tokenizer/generation contract, so no output-text comparison is made. Withheld does not mean unchanged.
- Overall relative weight change 100.4% over 338 parameter tensors (relative to the reference norm, so values above 100% are possible); the per-block table is in the report.
- Recorded relation: combined artifact change β configuration, tokenizer and generation-settings semantics all differ. The report does not attribute the difference to fine-tuning or to weights alone, does not locate edited weights and does not identify a cause. A is not presented as the immediate training parent of B; the comparison describes the difference between the two named artifacts only.
This is not a quality, safety or deployment grade and it does not rank models.
Report: https://tetracta-model-xray-sample-reports.static.hf.space/current/09-mukayese-qwen25-1.5b-math.html (it links to a limited Tetracta operational receipt; the receipt is a Tetracta record, not independent proof). Scope and limits, including what the instrument does not claim: https://www.tetracta.ai/model-xray/scope/ Β· Correction record: https://www.tetracta.ai/model-xray/correction/
Both repositories are on the eligible list of the free beta; a registered account can repeat this pair from the exact revisions above (the free allowance is 20 browser scans and five distinct source models per calendar month, so one pair uses two of the five). The comparable fields are the categorical ones listed here; numerical profiles are not published, and no cross-device bitwise equality is claimed.
If a different reference checkpoint is the appropriate base for this model, we would like to know which one, so that the recorded pair can be re-run against it. Corrections welcome. β Tetracta
Correction to one clause in the note above, posted the same day. It says one pair uses two of the five distinct source models in the monthly allowance. It does not: the allowance is counted on the source (A-side) repository of a job, and a comparison is a single job, so one pair uses one of the five. The rest of the note is unchanged. Scope and limits: https://www.tetracta.ai/model-xray/scope/ Correction record: https://www.tetracta.ai/model-xray/correction/ - Tetracta