RedHatAI/Mistral-Small-4-119B-2603.dflash

This is a DFlash speculator model for mistralai/Mistral-Small-4-119B-2603.

Training Details

This model was trained using the Speculators library on a subset of Magpie-Align/Magpie-Llama-3.1-Pro-300K-Filtered and the train_sft split of HuggingFaceH4/ultrachat_200k. Responses were regenerated by Mistral-Small-4-119B-2603. The model was initialized from a full-attention DFlash checkpoint and finetuned with sliding window attention (window=2048) on all 5 draft layers.

Commands

Using the Speculators library and the helper scripts provided in the repo.

Prepare data

# In virtual environment with speculators installed
python scripts/prepare_data.py \
  --model mistralai/Mistral-Small-4-119B-2603 \
  --data ./regenerated_data.jsonl \
  --output ./output/dflash_mistral_119b_sliding_window \
  --seq-length 8192

Launch vLLM

# In (separate) virtual environment with vllm installed
CUDA_VISIBLE_DEVICES=0,1,2,3 python scripts/launch_vllm.py \
  mistralai/Mistral-Small-4-119B-2603 \
  --target-layer-ids 2 18 33 \
  -- --port 8000 \
  --tensor-parallel-size 4 \
  --max-model-len 8194 \
  --enforce-eager

Launch training

Must be run once vLLM has finished launching and is running in the background.

# In virtual environment with speculators installed
CUDA_VISIBLE_DEVICES=4,5,6,7 torchrun \
  --standalone \
  --nproc_per_node 4 \
  scripts/train.py \
  --verifier-name-or-path mistralai/Mistral-Small-4-119B-2603 \
  --speculator-type dflash \
  --from-pretrained ./output/dflash_mistral_119b_sliding_window/patched_checkpoint \
  --data-path ./output/dflash_mistral_119b_sliding_window \
  --vllm-endpoint http://localhost:8000/v1 \
  --save-path ./output/dflash_mistral_119b_sliding_window/checkpoints \
  --epochs 3 \
  --lr 5e-5 \
  --total-seq-len 8192 \
  --max-anchors 1024 \
  --loss-fn '{"ce": 0.1, "tv": 0.9}' \
  --fsdp-shard \
  --no-sample-from-anchor \
  --on-missing generate \
  --on-generate delete \
  --checkpoint-freq 0.1

Model Specifications

Base Model mistralai/Mistral-Small-4-119B-2603
Chat Template mistralai/Mistral-Small-4-119B-2603 (use /chat/completions endpoint)
Format Safetensors
License Apache 2.0
Validation Hardware Nvidia H100

Deployment

# Install vLLM from the required PR
pip install git+https://github.com/vllm-project/vllm.git@refs/pull/41880/head

# Deploy with speculative decoding
vllm serve mistralai/Mistral-Small-4-119B-2603 \
    --tensor-parallel-size 4 \
    --max-model-len 16384 \
    --speculative-config '{
        "model": "RedHatAI/Mistral-Small-4-119B-2603.dflash",
        "num_speculative_tokens": 7,
        "method": "dflash"
    }'

Preliminary Evaluations

Per-position token acceptance rates across datasets:

Dataset Pos 1 Pos 2 Pos 3 Pos 4 Pos 5 Pos 6 Pos 7 Avg Length
HumanEval 80.9% 61.0% 45.7% 33.7% 24.7% 17.8% 12.1% 3.759
math_reasoning 91.7% 77.5% 65.5% 55.0% 45.1% 36.0% 27.4% 4.982
qa 62.6% 35.5% 19.5% 11.1% 6.0% 3.1% 1.4% 2.392
question 67.5% 42.3% 27.1% 17.9% 12.0% 8.0% 5.2% 2.801
rag 65.9% 39.7% 22.8% 13.5% 8.1% 4.2% 2.2% 2.565
summarization 65.1% 36.2% 19.3% 10.7% 5.6% 2.8% 1.0% 2.407
tool_call 66.6% 40.7% 24.9% 15.8% 10.0% 6.1% 3.5% 2.675
translation 71.5% 40.9% 20.1% 9.0% 4.0% 1.5% 0.8% 2.478
writing 67.9% 42.2% 26.6% 17.2% 11.8% 7.8% 5.0% 2.786

Latency Speedup

Speedup comparisons of DFlash speculative decoding vs. Eagle3 at varying request rates on Nvidia H100: image

References

Paper: DFlash: Block Diffusion for Flash Speculative Decoding

Downloads last month
147
Safetensors
Model size
2B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for RedHatAI/Mistral-Small-4-119B-2603.dflash

Finetuned
(13)
this model

Collection including RedHatAI/Mistral-Small-4-119B-2603.dflash

Paper for RedHatAI/Mistral-Small-4-119B-2603.dflash