Join the conversation

Join the community of Machine Learners and AI enthusiasts.

Sign Up
kvyb 
posted an update 9 days ago
Post
117
Qwen3.8-27B-Humanlike-Chat 2.0 is out.

It still texts like a person with no system prompt. Now it also calls tools, follows instructions, and writes the formal email you asked for, then goes back to texting.

vs the abliterated base it's built on:
- IFBench 37.3 → 43.7
- When2Call 48 → 58
- A blind judge took its reply for the real person's 23.5% of the time (official Qwen3.8-27B: 15.1%)

Trained with on-policy distillation: the model writes its own replies and a teacher grades every token.

GGUF from IQ4_XS to BF16, plus a free API, no key needed.

👉 LessThanThreeAI/Qwen3.8-27B-Humanlike-Chat-GGUF
Try it: LessThanThreeAI/Qwen3.8-27B-Humanlike-Chat

BFCL irrelevance is the one gain on this card that survives any pairing of the items.

I only used your Results table, so this is marginals only.

60 to 78 on 100 items means 2.0 fixed 18 more items than it broke. The base got 40 wrong, so at most 62 items can disagree between the two. Worst case McNemar z = 18/sqrt(62) = 2.29, exact sign test p = 0.030. It holds whichever items flipped.

The rest hang on a number the card doesn't print: how many items changed answer.

  • IFBench, 112 to 131 of 300: clears z = 1.96 only if 93 or fewer items flipped. Worst case z = 1.22.
  • When2Call, 48 to 58 of 100: needs 26 or fewer flips. Worst case 1.03.
  • MMLU-Pro, 157 to 145 of 200: needs 37 or fewer. Worst case 1.21.

That last one is the part I'd look at. Unpaired, the IFBench gain is 1.60 SE and the MMLU-Pro drop is 1.40 SE. Same yardstick, about the same strength. If IFBench is the cleanest win, MMLU-Pro is nearly as clean a loss.

Small metadata note while I was in the card: base_model_relation: quantized files 2.0 as one of the 82 quants of the Huihui base. The Hub lists 3 fine-tunes of that base and 2.0 isn't one of them, though it carries two merged LoRAs. finetune would put it there.

Every eval is greedy on the same prompts, so the flip counts should fall straight out of your per-item logs. How many IFBench and MMLU-Pro items changed answer between the base and 2.0?