Bruv1-4B

Bruv1-4B is a locally fine-tuned Qwen3.5-4B text decision model. It reads an evidence state, a question, and two to sixteen labeled options, then scores the next-token letters A–P. The merged checkpoint works with Hugging Face Transformers. The accompanying .kevala pack runs in Kevala. Training code and pinned recipes are open source.

The starting data mixture and one-epoch LoRA settings come from Together's Tev1 recipe at revision 57399714e6b5ef215c9c821cab35a7d5fbc5482b. Bruv is a separate local training implementation and does not use Together's service or weights. It uses Kevala's direct-option prompt and conditional label loss rather than Tev1's full completion and EOS loss.

Training

  • Base: Qwen/Qwen3.5-4B revision 851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a.
  • Data: all 37,840 Tev1 new-v1 training records, one epoch; 4,568 development records. The sources are MultiNLI, BoolQ, Banking77, AG News, SST-5, synthetic policies, routing, and research classification. Records with 24 options were deterministically reduced to 16 while retaining the correct answer. No raw records are redistributed here.
  • Objective: cross-entropy over only the valid A–P next-token logits at the end of Kevala's non-thinking Qwen chat prompt. No sequence packing.
  • Optimizer: LoRA rank 8, alpha 16, all linear transformer projections, AdamW, learning rate 5e-5, 3% warmup, cosine decay, effective batch 8 (1 × 8 accumulation), seed 42, maximum input 2,048 tokens. All 37,840 records fit without truncation.
  • Hardware: NVIDIA GeForce RTX 5070 Ti 16 GB, BF16, gradient checkpointing. 4,730 optimizer steps and full development scoring took 8,900 seconds. Peak PyTorch allocated GPU memory was 8.70 GiB.
  • Training input SHA-256: 81eba531a7e49b7846cf94f386cca6b493de103f5a071676579e985db77697a5.

Tev1's table and builder total 37,840 training examples. The article's prose also states 38,340; that figure does not match its table or this pinned build. Tev1's published starting script specifies one epoch. Its exact historical hosted job configuration was not independently verified.

Evaluation

The adapter scored 4,188/4,568 (91.68%) on development records immediately after training. This is a conditional option-choice result, not a general text-generation score. The evaluation splits share task families and construction methods with training; they do not establish general reasoning ability, safety behavior, or calibrated confidence. Kevala's separate browser decision ranking measures the quantized pack on different fixtures and rotating option orders.

The exported BF16 checkpoint scored 4,194/4,568 (91.81%) on the same development records and 2,481/2,800 (88.61%) on Tev1's separate test split. The test summary records source-level counts and hashes for the checkpoint and evaluation file. The test was scored on two CPU threads with batch size one to stay within the workstation's power budget; its accuracy is not a CPU latency result.

The Q8 pack is 4,751,303,168 bytes, SHA-256 ccbf575a60d33c3cce2b20e76a015b095a0ea149cdc80e29f954ad1ee432c0eb. All 12 scoreable reference cases matched exact prompt tokens and answer choices after quantization; maximum absolute option-score difference was 0.019503. One one-option fixture was outside the direct-options contract. On Kevala's separate Firefox WebGPU suite, the pack answered 799/864 decisions correctly with no invalid answers: 100/108 Kevala-authored, 403/432 SemIf-authored, and 296/324 SemIf perturbation decisions. These fixtures are small synthetic tasks, not a general model ranking or safety evaluation. The frozen-base comparison and Tev transfer and research challenge splits remain in progress.

Use

from transformers import AutoModelForCausalLM, AutoTokenizer

repo = "bvolpato/bruv1-4b"
tokenizer = AutoTokenizer.from_pretrained(repo)
model = AutoModelForCausalLM.from_pretrained(repo, dtype="auto")

Use the Bruv evaluation and prompt code to score option letters consistently. Ordinary text generation is not the evaluated behavior.

The Qwen base weights are Apache 2.0, Bruv code is MIT, and the public source datasets have separate terms. See Tev1's source provenance. This model card does not assert that every underlying dataset permits unrestricted commercial redistribution.

Downloads last month
309
Safetensors
Model size
4B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for bvolpato/bruv1-4b

Finetuned
Qwen/Qwen3.5-4B
Finetuned
(868)
this model
Quantizations
1 model