Kitchoon-26B-A4B
Gemma 4 RP model: a fused OV LoRA plus a separately trained lm_head, both on top of Pantheon-Reasoning 1.1 (heretic).
Pantheon already writes clean, active prose with low slop. I wanted to keep that and add what it lacks: more decisive characters, livelier dialogue, and stronger ERP, without dragging the dataset's clichés along with it.
What was changed
| Base (Pantheon 1.1 heretic) | Kitchoon | |
|---|---|---|
Attention OV (v_proj + o_proj) |
stock | LoRA, layers 10–29, 36 targets, baked at 0.70 |
lm_head |
stock | trained separately, blended with the stock head at 0.40 |
| QK, MoE experts, router, MLP, embeddings, vision tower | stock | identical |
OV LoRA details
| Parameter | Value |
|---|---|
| Rank / alpha | 32 / 64 |
| Learning rate | 2e-5, cosine |
| Epochs | 1 (931 steps) |
| Context | 3328 tokens |
| Eval loss | 4.67 → 1.45 |
How the head was trained (and why slop went down, not up)
The head was trained as a full tensor with the body frozen.
- Slop was masked out of the loss. Stock cliché phrases plus ~400 n-grams that showed up across many different characters in the dataset ("voice dropping to a", "a flicker of", "tilts her head"...) got no gradient. Same for text copied verbatim from the character card.
- Rare tokens were frozen. Rows for tokens seen fewer than 5 times in the targets stayed bit-identical to the base.
I would like to express my immense gratitude to Indexnusrefather, with whom I work, for providing the dataset.
Dataset
English RP dataset, ~2.8k multi-turn conversations built on character cards.
Cleaned before training:
- removed cards with underage characters and real people;
- fixed ~640 system prompts that instructed the model to write for the user;
- long conversations split into chunks instead of being truncated;
- the first and/or last paragraph randomly trimmed in part of the replies, to break up identical openings and hook-question endings;
- character name prefixes (
Name:) stripped from replies, since Pantheon was trained without them.
Using it
Recommended settings:
Parameter Value Temperature 1.0Min-P 0.05Thinking off — the training data had no reasoning blocks, and thinking mode was not tested.
What to expect
In testing, compared to the base:
- characters stick closer to the card — they pull in side characters, habits and lore from the description on their own;
- more dialogue, more humor, less of the character explaining themselves;
- scenes carry forward from turn to turn instead of resetting.
Known quirks, still there:
- characters sometimes repeat the user's question back before answering;
- replies like to end on an "X — or Y?" question;
- smirk and the em dash have not fully left the building.
Feedback
This is an experiment, and I mostly test it on my own cards, so my coverage is narrow. If you try it, I would really like to know:
- How it holds up in long sessions — scene logic, who is where, who is talking to whom.
- Whether the quirks above bother you in actual play.
- Your sampler settings, if you got clearly better or worse results than with the defaults.
Open a thread in the Community tab, I read all of them.
Credits
Thanks to Gryphe for Pantheon-Reasoning, and to Vortex5 for the heretic version it is built on. Thanks to the entire 26B-Suite team for their intellectual support.
Special thanks to Naphula and redaihf. You guys are awesome!
And special thanks, of course, to Indexnusrefather for his excellent dataset.
- Downloads last month
- 30
Model tree for SubMaroon/Kitchoon-26B-A4B
Base model
google/gemma-4-26B-A4B