Instructions to use willhx/Qwen3-8B-SDFT-MLE-Math-Search-Ecom with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use willhx/Qwen3-8B-SDFT-MLE-Math-Search-Ecom with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="willhx/Qwen3-8B-SDFT-MLE-Math-Search-Ecom") messages = [ {"role": "user", "content": "Who are you?"}, ] pipe(messages)# pip install -U transformers accelerate # Load model directly from transformers import AutoTokenizer, AutoModelForCausalLM tokenizer = AutoTokenizer.from_pretrained("willhx/Qwen3-8B-SDFT-MLE-Math-Search-Ecom") model = AutoModelForCausalLM.from_pretrained("willhx/Qwen3-8B-SDFT-MLE-Math-Search-Ecom", device_map="auto") messages = [ {"role": "user", "content": "Who are you?"}, ] inputs = tokenizer.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=256) print(tokenizer.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use willhx/Qwen3-8B-SDFT-MLE-Math-Search-Ecom with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "willhx/Qwen3-8B-SDFT-MLE-Math-Search-Ecom" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "willhx/Qwen3-8B-SDFT-MLE-Math-Search-Ecom", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/willhx/Qwen3-8B-SDFT-MLE-Math-Search-Ecom
- SGLang
How to use willhx/Qwen3-8B-SDFT-MLE-Math-Search-Ecom with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "willhx/Qwen3-8B-SDFT-MLE-Math-Search-Ecom" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "willhx/Qwen3-8B-SDFT-MLE-Math-Search-Ecom", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "willhx/Qwen3-8B-SDFT-MLE-Math-Search-Ecom" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "willhx/Qwen3-8B-SDFT-MLE-Math-Search-Ecom", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Docker Model Runner
How to use willhx/Qwen3-8B-SDFT-MLE-Math-Search-Ecom with Docker Model Runner:
docker model run hf.co/willhx/Qwen3-8B-SDFT-MLE-Math-Search-Ecom
Qwen3-8B-SDFT-MLE-Math-Search-Ecom
A full dense Qwen3-8B checkpoint formed by a weighted average of three SDFT models. All 399 parameter tensors are merged, including embeddings, the language model head, normalization weights, and attention/MLP projections.
Merge recipe
| Source model | Weight | Source revision |
|---|---|---|
| willamazon1/Qwen3-8B-SDFT-Math-LoRA-new | 0.2 | 9cdb66ee7ba89e4da0bf35719a3c04dcfcbd465d |
| willamazon1/sdft-search-lora-iter160 | 0.4 | 843d6b195c403c64eb42d7cae074ec0ce1c6b41a |
| willamazon1/sdft-tau-lora-iter160 | 0.4 | fe981d4463ce388ff96c2e9c6d3a25290540601c |
merged = BF16(0.2 * FP32(math) + 0.4 * FP32(search) + 0.4 * FP32(tau))
Each source already contains merged LoRA updates. Their non-LoRA backbone weights also differ, so this operation averages the complete dense checkpoints. The configuration and tokenizer assets are identical across the three sources. The Ecom component in the model name refers to the tau-bench retail checkpoint.
Model format and validation
- Architecture:
Qwen3ForCausalLM, 36 layers, 8,190,735,360 parameters. - Storage: BF16, four safetensors shards, approximately 16.4 GB.
- Vocabulary size: 151,936.
- All 399 saved tensors were checked element by element against the FP32 merge formula after BF16 rounding; zero mismatches and zero NaN/Inf elements.
- Tensor keys, shapes, dtypes, shard index, config loading, and tokenizer encoding
were verified. See
merge_validation.jsonfor shard SHA256 hashes.
Numerical validation does not establish task performance. This merged checkpoint has not been benchmarked for math, search, or retail tasks.
Usage
import torch
from transformers import AutoModelForCausalLM, AutoTokenizer
model_id = "willhx/Qwen3-8B-SDFT-MLE-Math-Search-Ecom"
tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(
model_id, torch_dtype=torch.bfloat16, device_map="auto"
)
inputs = tokenizer("Question: What is 12 * 8?\nAnswer:", return_tensors="pt").to(model.device)
outputs = model.generate(**inputs, max_new_tokens=64, do_sample=False)
print(tokenizer.decode(outputs[0], skip_special_tokens=True))
Use the prompt format expected by your evaluation or agent setup. The source models derive from Qwen3-8B-Base through SFT/RL stages; a general chat interface has not been validated for this merge. No separate PEFT adapter is needed.
- Downloads last month
- 423