Text Generation
Transformers
Safetensors
qwen3
Merge
taskvector
code
think
conversational
text-generation-inference
Instructions to use Montalte/qwen4b-code-think-taskvector with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use Montalte/qwen4b-code-think-taskvector with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="Montalte/qwen4b-code-think-taskvector") messages = [ {"role": "user", "content": "Who are you?"}, ] pipe(messages)# pip install -U transformers accelerate # Load model directly from transformers import AutoTokenizer, AutoModelForCausalLM tokenizer = AutoTokenizer.from_pretrained("Montalte/qwen4b-code-think-taskvector") model = AutoModelForCausalLM.from_pretrained("Montalte/qwen4b-code-think-taskvector", device_map="auto") messages = [ {"role": "user", "content": "Who are you?"}, ] inputs = tokenizer.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=256) print(tokenizer.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use Montalte/qwen4b-code-think-taskvector with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "Montalte/qwen4b-code-think-taskvector" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Montalte/qwen4b-code-think-taskvector", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/Montalte/qwen4b-code-think-taskvector
- SGLang
How to use Montalte/qwen4b-code-think-taskvector with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "Montalte/qwen4b-code-think-taskvector" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Montalte/qwen4b-code-think-taskvector", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "Montalte/qwen4b-code-think-taskvector" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Montalte/qwen4b-code-think-taskvector", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Docker Model Runner
How to use Montalte/qwen4b-code-think-taskvector with Docker Model Runner:
docker model run hf.co/Montalte/qwen4b-code-think-taskvector
Download MANIFEST.sha256 from Montalte/qwen4b-code-think-taskvector: direct link, hf CLI and curl.
- Browser
- Download file 1.82 kB
-
https://huggingface.co/Montalte/qwen4b-code-think-taskvector/resolve/main/MANIFEST.sha256
- Command line
-
hf download hf://Montalte/qwen4b-code-think-taskvector/MANIFEST.sha256
-
curl -L -o MANIFEST.sha256 https://huggingface.co/Montalte/qwen4b-code-think-taskvector/resolve/main/MANIFEST.sha256
1.82 kB
| 832dd9e00a68dd83b3c3fb9f5588dad7dcf337a0db50f7d9483f310cd292e92e LICENSE | |
| 5e66b33663c05d41fe3758efd5dde9aaa83e587e032e6f4d458a6fc74bc881f9 OFFICIAL_MERGE_RECEIPT.json | |
| 1592a771cd3fe8b9e6a976c36a2f8edbcd8c373b4ff5645c69911f0944802f33 README.md | |
| 490367bf1e2bd55f5a7381dc4083bcffc18430bcbc26457680c5c17268225bcc adapter/MANIFEST.json | |
| 16a5f02dafeca4c7804f7306e353375ac567d8158ed37baa20b80dcfe9188f33 adapter/TOKEN_ROWS_META.json | |
| 4e04ab4bb3a660f837b6f57d3964cebfdf5b5a58cebc0c5cd9a30372a396212c adapter/adapter_config.json | |
| 619dcee8c5c04416fed0cee928375583332ce2986bf3f51c0caf8408433831ed adapter/adapter_model.safetensors | |
| 407274bf3280d79ed8efabaf3c3b3e75debaeb6b8efdb4433ac0bae934144688 adapter/token_rows_both_sides.safetensors | |
| 87a2728cb8dc9fe424d624542f6060ec05a1d285ebbec578bb078900e33396b5 chat_template.jinja | |
| 40bcf4e0a6a7b5eb8aa582ee23f0c9700d76f147d17159419893bbad0f762176 config.json | |
| a2ac0d07aac37502207a212b122abe32bd9564266594a317c0ff490ebce1f226 generation_config.json | |
| 0e43b8415e406761c7ab73c8225fde51abdb7f51a4b80a9e90c616f53c3ea419 model.safetensors | |
| 5ebf618c3d547252a8078665415fa4fb3b080d071c7e0f2750f6af44e231bc3e provenance/DEV256_COMPLETE.json | |
| 24fc75c1ce08f0e4b7dcc31684c36c1f95b7032999c0e40308426781728f939c provenance/MQ0_NEAR_DUP.json | |
| c9f6e3b33a4693354da05fd6f18b3d05932430ca7c8dc9f7f008d3dcc49ac592 provenance/POLICY.json | |
| 5c45fca93b73d0293c39913fc74160cc24f3ed5efa4bfd51582428cc029e1cd2 provenance/RUN_IDENTITY.json | |
| 0304abe092d416ca7974b8f98f39783c52e71467bea8382ed000315ed917ec8a provenance/TRAINING_CONFIG.json | |
| 2ff0f61e200d18dd55b1301f9573e923abd6fe87608220f7126cf709d4cb7288 provenance/build_mix_distill_payload.py | |
| be75606093db2094d7cd20f3c2f385c212750648bd6ea4fb2bf507a6a4c55506 tokenizer.json | |
| 1749ac66e3e7ce3862337c23cca5895a2ee131929563e18a944f8c1c80363712 tokenizer_config.json | |