Elixir Gemma 4 12B

A Gemma 4 12B model specialized for writing Elixir. It starts from the instruction-tuned QAT checkpoint and is aimed at code that parses, compiles, and passes tests.

The repository contains the merged BF16 weights and a Q4_K_M GGUF for llama.cpp and LM Studio.

Elixir coding results

Scored on a 128-task executable harness. Each task asks for one Elixir module. A pass means the completion parses, compiles, and its tests pass on Elixir 1.20.2 / OTP 29. One completion per task, temperature 0, 16,384-token context, llama.cpp. The base and the fine-tune were both scored as 4-bit GGUF under that same setup.

Gemma 4 12B QAT Elixir Gemma 4 12B Change
Tasks passed 31 / 128 46 / 128 +15
Pass rate 24.2% 35.9% +11.7 pp
Relative gain +48%

The fine-tune solves about half again as many Elixir tasks as the base model it was trained from.

Where the gain shows up

Area Base Elixir Gemma 4 12B
Standard library 6 / 16 10 / 16
Error handling 8 / 16 11 / 16
OTP 2 / 16 4 / 16
BEAM 1 / 16 3 / 16
Abstractions 1 / 16 3 / 16

Completions are also cleaner: warning-free rate moves from 46.9% to 58.6%, and every completion stays inside a single fenced answer.

Run it

llama.cpp

llama-server -m Elixir-Gemma-4-12B-Q4_K_M.gguf -c 16384 --temp 1.0 --top-k 64 --top-p 0.95

The benchmark file is Elixir-Gemma-4-12B-Q4_K_M.gguf.

Transformers

Needs a Transformers build that knows gemma4_unified.

import torch
from transformers import AutoTokenizer, Gemma4UnifiedForConditionalGeneration

model_id = "jmarceno/Elixir-Gemma-4-12B"
tokenizer = AutoTokenizer.from_pretrained(model_id)
model = Gemma4UnifiedForConditionalGeneration.from_pretrained(
    model_id,
    dtype=torch.bfloat16,
    device_map="auto",
)

messages = [
    {
        "role": "user",
        "content": "Write an Elixir module Counter with a GenServer that stores an integer and supports increment and get.",
    }
]
prompt = tokenizer.apply_chat_template(messages, tokenize=False, add_generation_prompt=True)
inputs = tokenizer(prompt, return_tensors="pt").to(model.device)
output = model.generate(**inputs, max_new_tokens=512)
print(tokenizer.decode(output[0][inputs["input_ids"].shape[-1]:], skip_special_tokens=True))

License

Apache 2.0, the same license as google/gemma-4-12B-it-qat-q4_0-unquantized.

Downloads last month
286
Safetensors
Model size
12B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for jmarceno/Elixir-Gemma-4-12B

Quantized
(77)
this model
Quantizations
1 model