Instructions to use LessThanThreeAI/Qwen3.8-27B-Humanlike-Chat-2.0 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use LessThanThreeAI/Qwen3.8-27B-Humanlike-Chat-2.0 with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="LessThanThreeAI/Qwen3.8-27B-Humanlike-Chat-2.0") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] pipe(text=messages)# Load model directly from transformers import AutoProcessor, AutoModelForMultimodalLM processor = AutoProcessor.from_pretrained("LessThanThreeAI/Qwen3.8-27B-Humanlike-Chat-2.0") model = AutoModelForMultimodalLM.from_pretrained("LessThanThreeAI/Qwen3.8-27B-Humanlike-Chat-2.0", device_map="auto") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] inputs = processor.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=40) print(processor.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use LessThanThreeAI/Qwen3.8-27B-Humanlike-Chat-2.0 with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "LessThanThreeAI/Qwen3.8-27B-Humanlike-Chat-2.0" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "LessThanThreeAI/Qwen3.8-27B-Humanlike-Chat-2.0", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/LessThanThreeAI/Qwen3.8-27B-Humanlike-Chat-2.0
- SGLang
How to use LessThanThreeAI/Qwen3.8-27B-Humanlike-Chat-2.0 with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "LessThanThreeAI/Qwen3.8-27B-Humanlike-Chat-2.0" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "LessThanThreeAI/Qwen3.8-27B-Humanlike-Chat-2.0", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "LessThanThreeAI/Qwen3.8-27B-Humanlike-Chat-2.0" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "LessThanThreeAI/Qwen3.8-27B-Humanlike-Chat-2.0", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Docker Model Runner
How to use LessThanThreeAI/Qwen3.8-27B-Humanlike-Chat-2.0 with Docker Model Runner:
docker model run hf.co/LessThanThreeAI/Qwen3.8-27B-Humanlike-Chat-2.0
Qwen3.8-27B-Humanlike-Chat 2.0 (BF16 safetensors)
A 27B model that texts like a person and still does the work. No system prompt needed. This repo has the full BF16 weights for vLLM, SGLang and transformers.
Results, examples and how it was made are on the main card: Qwen3.8-27B-Humanlike-Chat-GGUF.
"Base" here always means huihui-ai/Huihui-Qwen3.8-27B-abliterated, an abliterated Qwen3.8-27B. It is not the official Qwen release.
vLLM
Tested with vLLM 0.27.1:
vllm serve LessThanThreeAI/Qwen3.8-27B-Humanlike-Chat-2.0 --served-model-name humanlike-2.0 \
--max-model-len 16384 --language-model-only \
--enable-auto-tool-choice --tool-call-parser qwen3_coder --reasoning-parser qwen3
Thinking is on by default. Send chat_template_kwargs: {"enable_thinking": false} per request for the fastest replies. Sampling: temperature 1.0, top_p 0.95, top_k 20 (the defaults in generation_config.json). For SGLang, point --model-path at this repo.
--language-model-only skips the vision tower. The weights keep the base's vision tower and MTP head unchanged, but 2.0 was tuned and tested on text only.
VRAM
The weights take 50.2 GiB in vLLM. Use one 80 GB card (H100, A100 80GB) or two 48 GB cards with --tensor-parallel-size 2. With 24 to 48 GB, take the FP8 repo or a GGUF from the main card.
Checks
Same check as every file on the main card: 5 fresh test chats against the reference (the base plus the 2.0 adapter at runtime in BF16), served on vLLM 0.27.1, one H100 NVL.
| KL vs reference | Top-1 agreement | Greedy turns identical | Check fails (greedy / sampled) | Decode tok/s |
|---|---|---|---|---|
| 0.0032 | 97.2% | 12 of 15 | 0 / 0 | 46 |
Live requests on the same server: "yo you up" with thinking off came back yeah what's up with no reasoning. With thinking on, the train question came back 17:35 with the reasoning in its own field. A weather request produced get_weather with {"city": "Lisbon", "unit": "celsius"}.
Commissions
Commissions open, DM codebottle on Discord (https://discord.com/users/320486798859960322). I build custom finetunes like this one: characters, product voices, distillation into smaller models, and domain or use-case specific models.
License
Apache-2.0, inherited from the upstream Qwen and Huihui releases.
- Downloads last month
- -