Instructions to use microsoft/Phi-4-mini-instruct with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use microsoft/Phi-4-mini-instruct with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="microsoft/Phi-4-mini-instruct", trust_remote_code=True) messages = [ {"role": "user", "content": "Who are you?"}, ] pipe(messages)# Load model directly from transformers import AutoTokenizer, AutoModelForCausalLM tokenizer = AutoTokenizer.from_pretrained("microsoft/Phi-4-mini-instruct", trust_remote_code=True) model = AutoModelForCausalLM.from_pretrained("microsoft/Phi-4-mini-instruct", trust_remote_code=True, device_map="auto") messages = [ {"role": "user", "content": "Who are you?"}, ] inputs = tokenizer.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=40) print(tokenizer.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Inference
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use microsoft/Phi-4-mini-instruct with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "microsoft/Phi-4-mini-instruct" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "microsoft/Phi-4-mini-instruct", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/microsoft/Phi-4-mini-instruct
- SGLang
How to use microsoft/Phi-4-mini-instruct with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "microsoft/Phi-4-mini-instruct" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "microsoft/Phi-4-mini-instruct", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "microsoft/Phi-4-mini-instruct" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "microsoft/Phi-4-mini-instruct", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Docker Model Runner
How to use microsoft/Phi-4-mini-instruct with Docker Model Runner:
docker model run hf.co/microsoft/Phi-4-mini-instruct
Up to 3.6x speedup with a DSpark speculative decoding draft model for Phi-4-mini-instruct (134 --> 486 tokens/sec; 4.95 accept length on math)
Phi-4-mini-instruct-DSpark
Model link: https://huggingface.co/rasyosef/Phi-4-mini-instruct-DSpark
A DSpark draft model for speculative decoding with microsoft/Phi-4-mini-instruct as the verifier, trained with speculators. The drafter proposes 8 tokens at a time and the verifier checks them in one forward pass, so output is identical to running the verifier alone β a lossless speedup. Mean acceptance length is 3.23 tokens committed per verification step, up to 4.95 on math_reasoning. That gives a throughput speedup over the verifier alone of 3.61Γ on math_reasoning and 3.44Γ on HumanEval, and 2.20Γ averaged across nine task types.
Training code: rasyosef/train-dspark-draft-models.
Trained on 116,000 samples.
Usage
vLLM loads the verifier automatically from the config β don't pass it separately.
vllm serve rasyosef/Phi-4-mini-instruct-DSpark \
--port 8000 \
--gpu-memory-utilization 0.8 \
--override-generation-config '{"temperature": 0}'
Then query the OpenAI-compatible endpoint at http://localhost:8000/v1. The override sets the server's default temperature to 0, matching the evaluation below; a request that passes its own temperature still takes precedence.
Details
4 Qwen3 layers (hidden size 3072, intermediate size 8192, 24 attention heads over 8 KV heads, head dim 128). The first three layers use sliding-window attention with a 2048-token window and the last layer uses full attention. Block size 8, aux hidden-state layers 2/10/22/30. The drafter predicts over a 50,000-token draft vocabulary, a subset of the verifier's 200,064 tokens.
Trained for 3 epochs at lr 3e-4 (cosine schedule with 4% warmup) on 116,000 Open PerfectBlend prompts regenerated by the verifier itself, split 96/4 into train and validation, with a {"ce": 0.1, "tv": 0.9} loss. Prompts prepared at 2048 tokens; training sequence length 4096, up to 512 anchors per sample. Verifier hidden states were pulled on demand from a running vLLM server during training and deleted after use rather than staged to disk up front.
Evaluation
evaluate.py throughput across the nine RedHatAI/speculator_benchmarks subsets, at temperature 0. acceptance_length is mean tokens committed per verification step, including the bonus token β floor 1.0, ceiling 9.0 at block size 8.
| subset | acceptance_length | pos_0 | pos_1 | pos_2 | pos_3 | pos_4 | pos_5 | pos_6 | pos_7 |
|---|---|---|---|---|---|---|---|---|---|
| math_reasoning | 4.954 | 86.8% | 74.1% | 62.3% | 51.1% | 41.3% | 33.1% | 26.5% | 20.2% |
| HumanEval | 4.736 | 85.1% | 70.8% | 58.6% | 47.3% | 38.6% | 30.2% | 24.0% | 18.9% |
| rag | 3.157 | 70.2% | 49.2% | 34.3% | 25.6% | 19.4% | 8.8% | 4.8% | 3.2% |
| writing | 3.073 | 68.7% | 45.9% | 31.3% | 22.0% | 15.8% | 11.6% | 7.5% | 4.6% |
| tool_call | 2.829 | 69.0% | 44.8% | 28.8% | 18.0% | 10.5% | 6.1% | 3.5% | 2.1% |
| qa | 2.646 | 63.2% | 39.6% | 25.0% | 17.8% | 14.2% | 2.7% | 1.5% | 0.7% |
| question | 2.618 | 65.1% | 39.2% | 23.2% | 14.0% | 8.6% | 5.5% | 3.7% | 2.4% |
| summarization | 2.327 | 62.8% | 35.5% | 19.0% | 9.2% | 4.1% | 1.6% | 0.4% | 0.1% |
| translation | 2.151 | 60.4% | 32.9% | 13.0% | 5.3% | 1.9% | 0.8% | 0.5% | 0.3% |
Weighted across all subsets: 3.234 over 69,478 verification steps.
Acceptance is highest where the verifier's next token is most predictable β math and code. math_reasoning leads HumanEval by about 0.2 tokens and holds a margin at every position in the block, and the two sit well clear of everything else: the next subset, rag, is more than 1.5 tokens behind HumanEval. math_reasoning's pos_4 (41.3%) is above pos_1 on qa, question, summarization, and translation (39.6% or lower). rag and writing follow at 3.16 and 3.07, tool_call at 2.83, and qa and question at 2.65 and 2.62. summarization (2.33) and translation (2.15) are lowest; on translation fewer than one draft in seven survives to pos_2. qa and rag also drop sharply between pos_4 and pos_5 (14.2% β 2.7% and 19.4% β 8.8%) rather than decaying smoothly.
Throughput (tokens/s)
Mean output throughput (tokens/s) across the same nine subsets, measured on a single A100 at concurrency 1, with DSpark at temperature 0. Baseline is Phi-4-mini-instruct with no speculative decoding, measured on every subset; without a drafter, throughput varies only modestly with the task (123.4β135.4 tokens/s). Speedup is the mean throughput relative to that subset's baseline.
| subset | baseline (no drafter) | DSpark (this model) | speedup |
|---|---|---|---|
| math_reasoning | 134.6 | 486.0 | 3.61Γ |
| HumanEval | 133.5 | 459.3 | 3.44Γ |
| rag | 123.4 | 219.1 | 1.78Γ |
| writing | 131.8 | 240.5 | 1.82Γ |
| tool_call | 123.8 | 287.3 | 2.32Γ |
| qa | 135.4 | 225.9 | 1.67Γ |
| question | 132.8 | 240.5 | 1.81Γ |
| summarization | 124.6 | 228.2 | 1.83Γ |
| translation | 127.9 | 190.6 | 1.49Γ |
| average speedup | 1.00Γ | β | 2.20Γ |
DSpark averages a 2.20Γ speedup (unweighted mean of per-subset speedups). The biggest gains are on math_reasoning at 3.61Γ (134.6 β 486.0 tokens/s) and HumanEval at 3.44Γ. tool_call is next at 2.32Γ; summarization, writing, question, and rag sit in a tight band between 1.78Γ and 1.83Γ, with qa at 1.67Γ. translation is lowest at 1.49Γ. Outside math and code, speedup does not follow acceptance length closely: tool_call is third on throughput but fifth on acceptance, and rag has the third-highest acceptance but one of the smaller speedups.
Means and medians agree to within about 6% on every DSpark subset. The mean runs about 5% above the median on tool_call, writing, and question, and 5β6% below it on math_reasoning and translation. Measured on medians instead, the speedup is 3.72Γ on math_reasoning, 3.51Γ on HumanEval, 2.06Γ on tool_call, and 2.15Γ on average; translation falls to 1.42Γ, because its baseline median (143.0 tokens/s) is well above its baseline mean (127.9).
Limitations
Works only with Phi-4-mini-instruct and is not usable as a standalone model. Acceptance falls off steeply past the first few positions on prose-like traffic, and on summarization and translation in particular a block size of 8 is mostly wasted β the gains concentrate in math and code. Translation is the weakest subset on both acceptance and throughput; the 50,000-token draft vocabulary, which covers a quarter of the verifier's multilingual vocabulary, may contribute to that. Throughput was measured at concurrency 1 on one A100 at temperature 0; real-world speedup depends on your traffic mix, hardware, sampling settings, and load, and because verification is lossless, the verifier's own behavior and biases carry through unchanged.
License
MIT, matching the verifier. The speculators training code is Apache-2.0.