gemma4-e4b-arcade-GGUF

Gemma 4 E4B, LoRA fine-tuned for ARCade (ARC aid via AI): run the Griffin nanopore preprocessing pipeline (Dorado basecalling, demultiplexing, FASTQ, FastQC/NanoPlot/MultiQC) on the University of Calgary ARC cluster by chatting.

ARCade downloads this file itself; you do not need to fetch it by hand.

Files

File Size MD5
gemma4-e4b-arcade-Q6_K.gguf 6,172,078,912 B 60305388062e2afea94e06ef52c66bde

Prompt format

ARCade renders the chat template itself and calls llama.cpp's /completion endpoint. Tool calls come out as

<|tool_call>call:TOOL_NAME{{"arg": "value"}}<tool_call|>

Training

  • Base: google/gemma-4-e4b-it, snapshot fee6332c1abaafb77f6f9624236c63aa2f1d0187.
  • LoRA with mlx-lm 0.31.3: rank 16, 16 layers, 640 iterations, batch 4, lr 1e-4.
  • Data: 1,264 synthetic multi-turn conversations. They cover the one-command workflow, the four steps run one at a time, job status, results, resources, and the questions participants ask.
  • Merged, converted with convert_hf_to_gguf.py --outtype f16, quantized with llama-quantize Q6_K. Q6_K, not Q4_K_M: on 60 held-out turns Q6_K matched the f16 model (48 vs 47 exact), Q4_K_M dropped to 37.

License

Apache 2.0, inherited from google/gemma-4-e4b-it.

Downloads last month
97
GGUF
Model size
7B params
Architecture
gemma4
Hardware compatibility
Log In to add your hardware

6-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support