Text Generation
Transformers
Safetensors
English
Chinese
Russian
yue2
music-generation
orbitquant
quantization
4-bit precision
custom-code
8-bit precision
Instructions to use WaveCut/YuE2-3B-OrbitQuant-W4A4 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use WaveCut/YuE2-3B-OrbitQuant-W4A4 with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="WaveCut/YuE2-3B-OrbitQuant-W4A4")# pip install -U transformers accelerate # Load model directly from transformers import AutoModelForCausalLM model = AutoModelForCausalLM.from_pretrained("WaveCut/YuE2-3B-OrbitQuant-W4A4", device_map="auto") - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use WaveCut/YuE2-3B-OrbitQuant-W4A4 with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "WaveCut/YuE2-3B-OrbitQuant-W4A4" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "WaveCut/YuE2-3B-OrbitQuant-W4A4", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker
docker model run hf.co/WaveCut/YuE2-3B-OrbitQuant-W4A4
- SGLang
How to use WaveCut/YuE2-3B-OrbitQuant-W4A4 with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "WaveCut/YuE2-3B-OrbitQuant-W4A4" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "WaveCut/YuE2-3B-OrbitQuant-W4A4", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "WaveCut/YuE2-3B-OrbitQuant-W4A4" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "WaveCut/YuE2-3B-OrbitQuant-W4A4", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }' - Docker Model Runner
How to use WaveCut/YuE2-3B-OrbitQuant-W4A4 with Docker Model Runner:
docker model run hf.co/WaveCut/YuE2-3B-OrbitQuant-W4A4
Download src/kernel-source/orbitquant-native/torch-ext/torch_binding.h from WaveCut/YuE2-3B-OrbitQuant-W4A4: direct link, hf CLI and curl.
- Browser
- Download file 2.45 kB
-
https://huggingface.co/WaveCut/YuE2-3B-OrbitQuant-W4A4/resolve/main/src/kernel-source/orbitquant-native/torch-ext/torch_binding.h
- Command line
-
hf download hf://WaveCut/YuE2-3B-OrbitQuant-W4A4/src/kernel-source/orbitquant-native/torch-ext/torch_binding.h
-
curl -L -o torch_binding.h https://huggingface.co/WaveCut/YuE2-3B-OrbitQuant-W4A4/resolve/main/src/kernel-source/orbitquant-native/torch-ext/torch_binding.h
2.45 kB
| using OrbitQuantTensor = torch::stable::Tensor; | |
| using OrbitQuantTensor = torch::Tensor; | |
| void matmul_packed_weight( | |
| OrbitQuantTensor &out, | |
| OrbitQuantTensor const &x, | |
| OrbitQuantTensor const &packed_weight_indices, | |
| OrbitQuantTensor const &row_norms, | |
| OrbitQuantTensor const ¢roids, | |
| OrbitQuantTensor const &bias, | |
| bool has_bias, | |
| int64_t bits, | |
| int64_t out_features, | |
| int64_t in_features, | |
| int64_t block_m, | |
| int64_t block_n, | |
| int64_t block_k); | |
| void quantize_activations_cpu( | |
| OrbitQuantTensor &out, | |
| OrbitQuantTensor const &x, | |
| OrbitQuantTensor const &permutation, | |
| OrbitQuantTensor const &signs, | |
| OrbitQuantTensor const ¢roids, | |
| OrbitQuantTensor const &boundaries, | |
| double eps, | |
| double inv_sqrt_block, | |
| int64_t block_size); | |
| void matmul_packed_adaln_int4_cpu( | |
| OrbitQuantTensor &out, | |
| OrbitQuantTensor const &x, | |
| OrbitQuantTensor const &packed_weight, | |
| OrbitQuantTensor const &scales, | |
| OrbitQuantTensor const &bias, | |
| bool has_bias, | |
| int64_t out_features, | |
| int64_t in_features, | |
| int64_t group_size); | |
| void matmul_packed_w4a4_int8( | |
| torch::Tensor &out, | |
| torch::Tensor const &packed_activations, | |
| torch::Tensor const &packed_weight_indices, | |
| torch::Tensor const &token_norms, | |
| torch::Tensor const &row_norms, | |
| torch::Tensor const &activation_codes, | |
| torch::Tensor const &weight_codes, | |
| torch::Tensor const &bias, | |
| bool has_bias, | |
| double activation_scale, | |
| double weight_scale, | |
| int64_t out_features, | |
| int64_t in_features, | |
| int64_t tile_m, | |
| int64_t tile_n, | |
| bool async_packed, | |
| bool weight_k_major); | |
| void quantize_activations_packed_w4( | |
| torch::Tensor &packed_out, | |
| torch::Tensor &norms_out, | |
| torch::Tensor const &x, | |
| torch::Tensor const &permutation, | |
| torch::Tensor const &signs, | |
| torch::Tensor const &boundaries, | |
| double eps, | |
| double inv_sqrt_block, | |
| int64_t threads); | |
| void quantize_activations_int8( | |
| torch::Tensor &int8_out, | |
| torch::Tensor &norms_out, | |
| torch::Tensor const &x, | |
| torch::Tensor const &permutation, | |
| torch::Tensor const &signs, | |
| torch::Tensor const &boundaries, | |
| torch::Tensor const &codes, | |
| double eps, | |
| double inv_sqrt_block, | |
| int64_t threads); | |