Text Generation
Transformers
Safetensors
Chinese
English
qwen3_dspark
feature-extraction
speculative-decoding
dspark
glm-5.2
draft-model
custom_code
Instructions to use AlayaNeW/GLM-5.2-DSpark with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use AlayaNeW/GLM-5.2-DSpark with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="AlayaNeW/GLM-5.2-DSpark", trust_remote_code=True)# pip install -U transformers accelerate # Load model directly from transformers import AutoModel model = AutoModel.from_pretrained("AlayaNeW/GLM-5.2-DSpark", trust_remote_code=True, device_map="auto") - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use AlayaNeW/GLM-5.2-DSpark with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "AlayaNeW/GLM-5.2-DSpark" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "AlayaNeW/GLM-5.2-DSpark", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker
docker model run hf.co/AlayaNeW/GLM-5.2-DSpark
- SGLang
How to use AlayaNeW/GLM-5.2-DSpark with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "AlayaNeW/GLM-5.2-DSpark" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "AlayaNeW/GLM-5.2-DSpark", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "AlayaNeW/GLM-5.2-DSpark" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "AlayaNeW/GLM-5.2-DSpark", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }' - Docker Model Runner
How to use AlayaNeW/GLM-5.2-DSpark with Docker Model Runner:
docker model run hf.co/AlayaNeW/GLM-5.2-DSpark
Download sampling_utils.py from AlayaNeW/GLM-5.2-DSpark: direct link, hf CLI and curl.
- Browser
- Download file 1.8 kB
-
https://huggingface.co/AlayaNeW/GLM-5.2-DSpark/resolve/main/sampling_utils.py
- Command line
-
hf download hf://AlayaNeW/GLM-5.2-DSpark/sampling_utils.py
-
curl -L -o sampling_utils.py https://huggingface.co/AlayaNeW/GLM-5.2-DSpark/resolve/main/sampling_utils.py
1.8 kB
| from __future__ import annotations | |
| import torch | |
| def logits_to_probs(logits: torch.Tensor, temperature: float) -> torch.Tensor: | |
| if temperature < 1e-5: | |
| probs = torch.zeros_like(logits, dtype=torch.float32) | |
| probs.scatter_(-1, torch.argmax(logits, dim=-1, keepdim=True), 1.0) | |
| return probs | |
| return torch.softmax(logits.float() / temperature, dim=-1) | |
| def sample_from_probs(probs: torch.Tensor) -> torch.Tensor: | |
| bsz, seq_len, vocab_size = probs.shape | |
| flat = probs.reshape(-1, vocab_size) | |
| return torch.multinomial(flat, num_samples=1).reshape(bsz, seq_len) | |
| def sample_tokens(logits: torch.Tensor, temperature: float = 0.0) -> torch.Tensor: | |
| if temperature < 1e-5: | |
| return torch.argmax(logits, dim=-1) | |
| bsz, seq_len, vocab_size = logits.shape | |
| flat_logits = logits.reshape(-1, vocab_size) / temperature | |
| probs = torch.softmax(flat_logits, dim=-1) | |
| return torch.multinomial(probs, num_samples=1).reshape(bsz, seq_len) | |
| def gather_token_probs(probs: torch.Tensor, token_ids: torch.Tensor) -> torch.Tensor: | |
| return probs.gather(dim=-1, index=token_ids.unsqueeze(-1)).squeeze(-1) | |
| def sample_residual( | |
| target_probs: torch.Tensor, | |
| draft_probs: torch.Tensor, | |
| ) -> torch.Tensor: | |
| residual = torch.clamp(target_probs - draft_probs, min=0.0) | |
| residual_mass = residual.sum(dim=-1, keepdim=True) | |
| if torch.any(residual_mass <= 1e-8): | |
| residual = torch.where(residual_mass <= 1e-8, target_probs, residual) | |
| residual_mass = residual.sum(dim=-1, keepdim=True) | |
| residual = residual / residual_mass.clamp_min(1e-8) | |
| return sample_from_probs(residual.unsqueeze(1)).squeeze(1) | |
| __all__ = [ | |
| "gather_token_probs", | |
| "logits_to_probs", | |
| "sample_from_probs", | |
| "sample_residual", | |
| "sample_tokens", | |
| ] | |