Text Generation
Transformers
PyTorch
MLX
English
tree-attention
apple-m4
mps
structured-generation
parallel-decoding
constrained-decoding
apple-silicon
classification
json
Instructions to use epsilon3/Qwen-2.5-1B-RLCD-Fast with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use epsilon3/Qwen-2.5-1B-RLCD-Fast with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="epsilon3/Qwen-2.5-1B-RLCD-Fast")# pip install -U transformers accelerate # Load model directly from transformers import AutoModel model = AutoModel.from_pretrained("epsilon3/Qwen-2.5-1B-RLCD-Fast", device_map="auto") - MLX
How to use epsilon3/Qwen-2.5-1B-RLCD-Fast with MLX:
# Make sure mlx-lm is installed # pip install --upgrade mlx-lm # if on a CUDA device, also pip install mlx[cuda] # Generate text with mlx-lm from mlx_lm import load, generate model, tokenizer = load("epsilon3/Qwen-2.5-1B-RLCD-Fast") prompt = "Once upon a time in" text = generate(model, tokenizer, prompt=prompt, verbose=True) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- LM Studio
- vLLM
How to use epsilon3/Qwen-2.5-1B-RLCD-Fast with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "epsilon3/Qwen-2.5-1B-RLCD-Fast" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "epsilon3/Qwen-2.5-1B-RLCD-Fast", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker
docker model run hf.co/epsilon3/Qwen-2.5-1B-RLCD-Fast
- SGLang
How to use epsilon3/Qwen-2.5-1B-RLCD-Fast with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "epsilon3/Qwen-2.5-1B-RLCD-Fast" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "epsilon3/Qwen-2.5-1B-RLCD-Fast", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "epsilon3/Qwen-2.5-1B-RLCD-Fast" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "epsilon3/Qwen-2.5-1B-RLCD-Fast", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }' - MLX LM
How to use epsilon3/Qwen-2.5-1B-RLCD-Fast with MLX LM:
Generate or start a chat session
# Install MLX LM uv tool install mlx-lm # Generate some text mlx_lm.generate --model "epsilon3/Qwen-2.5-1B-RLCD-Fast" --prompt "Once upon a time"
- Docker Model Runner
How to use epsilon3/Qwen-2.5-1B-RLCD-Fast with Docker Model Runner:
docker model run hf.co/epsilon3/Qwen-2.5-1B-RLCD-Fast
- Atomic Chat
File size: 3,190 Bytes
f36843f | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 | {
"environment": {
"model": "Qwen/Qwen2.5-1.5B-Instruct",
"revision": "989aa7980e4cf806f80c7fef2b1adb7bc71aa306",
"torch": "2.11.0",
"transformers": "4.57.6",
"gpu": null,
"platform": "macOS-26.6.2-arm64-arm-64bit",
"dtype": "torch.float32",
"attention": "sdpa"
},
"arguments": {
"model": "Qwen/Qwen2.5-1.5B-Instruct",
"presets": [
"fintech_fraud"
],
"fields": [
4
],
"steps": 4,
"repeats": 8,
"warmup": 2,
"output": "mac_cpu_continuation.json"
},
"cases": [
{
"preset": "fintech_fraud",
"fields": 4,
"prefix_tokens": 291,
"suffix_tokens": 28,
"batch_padded_tokens": 32,
"correctness": {
"max_abs_logit_error": 6.008148193359375e-05,
"mean_abs_logit_error": 6.594165370188421e-06,
"full_vocab_argmax_agreement": 1.0,
"candidate_decisions_equal": true,
"candidate_collision_fields": []
},
"decode": {
"batch": {
"median_ms": 622.972208991996,
"min_ms": 586.4072090043919,
"p90_ms": 801.8490410031518,
"samples_ms": [
586.4072090043919,
600.5369170015911,
603.2724169926951,
680.7952079980168,
644.0831670042826,
614.7183339926414,
631.2260839913506,
801.8490410031518
]
},
"tree": {
"median_ms": 585.8095005023642,
"min_ms": 559.2341670126189,
"p90_ms": 626.3385829952313,
"samples_ms": [
568.0354160140269,
596.4209999947343,
559.2341670126189,
578.7606250087265,
626.3385829952313,
579.8298340087058,
591.7891669960227,
613.8133749918779
]
}
},
"end_to_end": {
"batch": {
"median_ms": 967.1460625031614,
"min_ms": 940.2594160055742,
"p90_ms": 1015.3512500110082,
"samples_ms": [
999.5129159942735,
940.2594160055742,
1015.3512500110082,
962.8962919960031,
971.3958330103196,
977.6666249963455,
943.482041999232,
960.0878339988412
]
},
"tree": {
"median_ms": 938.4278540019295,
"min_ms": 914.2270419979468,
"p90_ms": 957.8748749918304,
"samples_ms": [
928.7943340023048,
914.2270419979468,
957.8748749918304,
952.5329169991892,
935.7932499988237,
941.0624580050353,
943.823458001134,
917.1714999974938
]
}
},
"memory": {},
"decode_speedup": 1.0634382140572365,
"end_to_end_speedup": 1.0306024681372818,
"decode_speedup_ci95": [
1.0187897902697607,
1.1586212607706694
],
"end_to_end_speedup_ci95": [
1.0099125150404291,
1.0610482668353582
]
}
]
}
|