Text Generation
Transformers
TensorBoard
Safetensors
English
qwen3
byte-level
pretraining
symbolic
text-generation-inference
Instructions to use dotlabs/void.1 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use dotlabs/void.1 with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="dotlabs/void.1")# pip install -U transformers accelerate # Load model directly from transformers import AutoTokenizer, AutoModelForCausalLM tokenizer = AutoTokenizer.from_pretrained("dotlabs/void.1") model = AutoModelForCausalLM.from_pretrained("dotlabs/void.1", device_map="auto") - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use dotlabs/void.1 with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "dotlabs/void.1" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "dotlabs/void.1", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker
docker model run hf.co/dotlabs/void.1
- SGLang
How to use dotlabs/void.1 with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "dotlabs/void.1" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "dotlabs/void.1", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "dotlabs/void.1" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "dotlabs/void.1", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }' - Docker Model Runner
How to use dotlabs/void.1 with Docker Model Runner:
docker model run hf.co/dotlabs/void.1
Download benchmark_cache/performance.json from dotlabs/void.1: direct link, hf CLI and curl.
- Browser
- Download file 6.05 kB
-
https://huggingface.co/dotlabs/void.1/resolve/main/benchmark_cache/performance.json
- Command line
-
hf download hf://dotlabs/void.1/benchmark_cache/performance.json
-
curl -L -o performance.json https://huggingface.co/dotlabs/void.1/resolve/main/benchmark_cache/performance.json
6.05 kB
| { | |
| "cache_key": { | |
| "version": 1, | |
| "kind": "gpu", | |
| "gpu": { | |
| "name": "NVIDIA RTX PRO 6000 Blackwell Server Edition", | |
| "total_memory": 101975851008, | |
| "capability": [ | |
| 12, | |
| 0 | |
| ] | |
| }, | |
| "torch": "2.8.0+cu128", | |
| "cuda": "12.8", | |
| "python": "3.13.11", | |
| "packages": { | |
| "transformers": "4.57.1", | |
| "numpy": "2.2.6", | |
| "pyarrow": "21.0.0", | |
| "huggingface_hub": "0.36.0" | |
| }, | |
| "settings": { | |
| "context": 2048, | |
| "global_batch": 90, | |
| "microbatch_candidates": [ | |
| 8, | |
| 16, | |
| 32, | |
| 48, | |
| 64, | |
| 72, | |
| 76, | |
| 80, | |
| 86, | |
| 90 | |
| ], | |
| "batch_selection": "max_batch", | |
| "gradient_checkpointing": false, | |
| "compile_model": true, | |
| "autotune_microbatch": true, | |
| "attention_backend": "flash", | |
| "vram_fraction": 0.98 | |
| }, | |
| "source": "d52752a64023cf29dc386d43835cf3173c5318ee093c3b375a812733c74023a1", | |
| "requested_microbatch_ceiling": 90 | |
| }, | |
| "cache_reused": true, | |
| "attention": "PyTorch FlashAttention (forced; forward/backward verified)", | |
| "selected": { | |
| "microbatch": 90, | |
| "mode": "compiled", | |
| "microbatch_seconds": 0.9440444560000287, | |
| "peak_gb": 89.37790203094482, | |
| "fits_budget": true, | |
| "estimated_update_seconds": 0.9440444560000287, | |
| "input_tokens_per_second": 195245.04257031978 | |
| }, | |
| "measurements": [ | |
| { | |
| "microbatch": 8, | |
| "mode": "eager", | |
| "microbatch_seconds": 0.15785124100000303, | |
| "peak_gb": 14.764304637908936, | |
| "fits_budget": true, | |
| "estimated_update_seconds": 1.7943123250000212, | |
| "input_tokens_per_second": 103793.92582665654 | |
| }, | |
| { | |
| "microbatch": 16, | |
| "mode": "eager", | |
| "microbatch_seconds": 0.3454984619999948, | |
| "peak_gb": 28.084960460662842, | |
| "fits_budget": true, | |
| "estimated_update_seconds": 1.92592258199997, | |
| "input_tokens_per_second": 94842.67979172798 | |
| }, | |
| { | |
| "microbatch": 32, | |
| "mode": "eager", | |
| "microbatch_seconds": 0.7818738260000089, | |
| "peak_gb": 54.827155113220215, | |
| "fits_budget": true, | |
| "estimated_update_seconds": 2.1863095480000254, | |
| "input_tokens_per_second": 83819.15063620413 | |
| }, | |
| { | |
| "microbatch": 48, | |
| "mode": "eager", | |
| "microbatch_seconds": 1.171907979999986, | |
| "peak_gb": 81.63183498382568, | |
| "fits_budget": true, | |
| "estimated_update_seconds": 2.1970569614999818, | |
| "input_tokens_per_second": 83883.71926608194 | |
| }, | |
| { | |
| "microbatch": 8, | |
| "mode": "compiled", | |
| "microbatch_seconds": 0.0880895730000475, | |
| "peak_gb": 9.113307476043701, | |
| "fits_budget": true, | |
| "estimated_update_seconds": 1.0067526180005189, | |
| "input_tokens_per_second": 185992.50106469658 | |
| }, | |
| { | |
| "microbatch": 16, | |
| "mode": "compiled", | |
| "microbatch_seconds": 0.1701795650000122, | |
| "peak_gb": 16.87287473678589, | |
| "fits_budget": true, | |
| "estimated_update_seconds": 0.9559394930000451, | |
| "input_tokens_per_second": 192549.55787434086 | |
| }, | |
| { | |
| "microbatch": 32, | |
| "mode": "compiled", | |
| "microbatch_seconds": 0.34178767000003063, | |
| "peak_gb": 32.49641752243042, | |
| "fits_budget": true, | |
| "estimated_update_seconds": 0.950241911000063, | |
| "input_tokens_per_second": 191744.77534544803 | |
| }, | |
| { | |
| "microbatch": 48, | |
| "mode": "compiled", | |
| "microbatch_seconds": 0.5061955770000282, | |
| "peak_gb": 48.123157024383545, | |
| "fits_budget": true, | |
| "estimated_update_seconds": 0.9457143725000208, | |
| "input_tokens_per_second": 194201.61784620755 | |
| }, | |
| { | |
| "microbatch": 64, | |
| "mode": "compiled", | |
| "microbatch_seconds": 0.6595429280000076, | |
| "peak_gb": 63.757601737976074, | |
| "fits_budget": true, | |
| "estimated_update_seconds": 0.9208904890000156, | |
| "input_tokens_per_second": 198731.56762890573 | |
| }, | |
| { | |
| "microbatch": 72, | |
| "mode": "compiled", | |
| "microbatch_seconds": 0.741346276999991, | |
| "peak_gb": 71.65071392059326, | |
| "fits_budget": true, | |
| "estimated_update_seconds": 0.9243867199999727, | |
| "input_tokens_per_second": 198903.00197730918 | |
| }, | |
| { | |
| "microbatch": 76, | |
| "mode": "compiled", | |
| "microbatch_seconds": 0.7816785220000497, | |
| "peak_gb": 75.5242338180542, | |
| "fits_budget": true, | |
| "estimated_update_seconds": 0.9251490005000278, | |
| "input_tokens_per_second": 199120.22093398403 | |
| }, | |
| { | |
| "microbatch": 80, | |
| "mode": "compiled", | |
| "microbatch_seconds": 0.8258502979999776, | |
| "peak_gb": 79.40062236785889, | |
| "fits_budget": true, | |
| "estimated_update_seconds": 0.9283437199999867, | |
| "input_tokens_per_second": 198389.46646478592 | |
| }, | |
| { | |
| "microbatch": 86, | |
| "mode": "compiled", | |
| "microbatch_seconds": 0.8883623469999975, | |
| "peak_gb": 85.4687147140503, | |
| "fits_budget": true, | |
| "estimated_update_seconds": 0.9419139469999607, | |
| "input_tokens_per_second": 198261.44207347915 | |
| }, | |
| { | |
| "microbatch": 90, | |
| "mode": "compiled", | |
| "microbatch_seconds": 0.9299377820000245, | |
| "peak_gb": 89.3772840499878, | |
| "fits_budget": true, | |
| "estimated_update_seconds": 0.9299377820000245, | |
| "input_tokens_per_second": 198206.80863571487 | |
| } | |
| ], | |
| "failures": [ | |
| { | |
| "microbatch": 64, | |
| "mode": "eager", | |
| "reason": "oom" | |
| }, | |
| { | |
| "microbatch": 72, | |
| "mode": "eager", | |
| "reason": "oom" | |
| }, | |
| { | |
| "microbatch": 76, | |
| "mode": "eager", | |
| "reason": "oom" | |
| }, | |
| { | |
| "microbatch": 80, | |
| "mode": "eager", | |
| "reason": "oom" | |
| }, | |
| { | |
| "microbatch": 86, | |
| "mode": "eager", | |
| "reason": "oom" | |
| }, | |
| { | |
| "microbatch": 90, | |
| "mode": "eager", | |
| "reason": "oom" | |
| } | |
| ], | |
| "selection_policy": "max_batch", | |
| "budget_gb": 92.52170804977416, | |
| "reserve_gb": 2.1791439056396484, | |
| "note": "Forward/backward throughput; excludes optimizer, data, logging and checkpoint time.", | |
| "legacy_seed_verified": false | |
| } |