Qwen3-8B-SDFT-MLE-Math-Search-Ecom

A full dense Qwen3-8B checkpoint formed by a weighted average of three SDFT models. All 399 parameter tensors are merged, including embeddings, the language model head, normalization weights, and attention/MLP projections.

Merge recipe

Source model Weight Source revision
willamazon1/Qwen3-8B-SDFT-Math-LoRA-new 0.2 9cdb66ee7ba89e4da0bf35719a3c04dcfcbd465d
willamazon1/sdft-search-lora-iter160 0.4 843d6b195c403c64eb42d7cae074ec0ce1c6b41a
willamazon1/sdft-tau-lora-iter160 0.4 fe981d4463ce388ff96c2e9c6d3a25290540601c
merged = BF16(0.2 * FP32(math) + 0.4 * FP32(search) + 0.4 * FP32(tau))

Each source already contains merged LoRA updates. Their non-LoRA backbone weights also differ, so this operation averages the complete dense checkpoints. The configuration and tokenizer assets are identical across the three sources. The Ecom component in the model name refers to the tau-bench retail checkpoint.

Model format and validation

  • Architecture: Qwen3ForCausalLM, 36 layers, 8,190,735,360 parameters.
  • Storage: BF16, four safetensors shards, approximately 16.4 GB.
  • Vocabulary size: 151,936.
  • All 399 saved tensors were checked element by element against the FP32 merge formula after BF16 rounding; zero mismatches and zero NaN/Inf elements.
  • Tensor keys, shapes, dtypes, shard index, config loading, and tokenizer encoding were verified. See merge_validation.json for shard SHA256 hashes.

Numerical validation does not establish task performance. This merged checkpoint has not been benchmarked for math, search, or retail tasks.

Usage

import torch
from transformers import AutoModelForCausalLM, AutoTokenizer

model_id = "willhx/Qwen3-8B-SDFT-MLE-Math-Search-Ecom"
tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(
    model_id, torch_dtype=torch.bfloat16, device_map="auto"
)

inputs = tokenizer("Question: What is 12 * 8?\nAnswer:", return_tensors="pt").to(model.device)
outputs = model.generate(**inputs, max_new_tokens=64, do_sample=False)
print(tokenizer.decode(outputs[0], skip_special_tokens=True))

Use the prompt format expected by your evaluation or agent setup. The source models derive from Qwen3-8B-Base through SFT/RL stages; a general chat interface has not been validated for this merge. No separate PEFT adapter is needed.

Downloads last month
423
Safetensors
Model size
8B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for willhx/Qwen3-8B-SDFT-MLE-Math-Search-Ecom