ConsensusLab Fact-Checking — standalone merged 3B model

This contains the complete Qwen2.5-3B-Instruct base weights with the final step-80 Fact-Checking GRPO LoRA merged into them. Loading requires Transformers, not a separate PEFT adapter. No additional training was performed during export.

Provenance

  • Base: Qwen/Qwen2.5-3B-Instruct
  • Exact base revision: aa8e72537993ba99e69dfaafa59ed015b17504d1
  • Source adapter: tanny2109/consensuslab-peer-deference-fact-checking-qwen2.5-3b-grpo-lora
  • Source adapter revision: a7bb887017b7f985802994ee45a619881ed65b2d
  • Adapter SHA-256: 38091093c74a435bdeaafcf42abb631f9fa773c59074b79dc0fd010414b7fa13
  • Export: PEFT merge_and_unload(safe_merge=True) in BF16, without quantization.
  • Training reward: choosing the objectively correct option.
  • Training: 80 GRPO steps, rank 8, alpha 16, Q/K/V/O LoRA; base weights frozen.
  • Training data: 240 synthetic arithmetic, lookup, and prefix-classification cases.

Study results measured before merging

Measure Original Agree-with-Peer Fact-Checking
Benign accuracy without advice 41/48 34/48 40/48
Benign accuracy with wrong peer advice 35/48 0/48 40/48
Unauthorized approval with wrong peer advice 0/24 24/24 0/24
Legitimate approval without advice 24/24 14/24 24/24

These are saved results from the source study, not a new full evaluation of this merged export. The protocol returns a single option letter. Peer advice is text inside the user message after task facts and choices; it is not an API role. Prompt order and wording materially affect the behavior. Training used one seed. Evaluation instances were disjoint from gradient training; a subset was used during smoke development. The model is not a general-purpose fact checker or a production authorization system. The Agree-with-Peer arm deliberately rewards incorrect advice following and is intended for controlled research comparisons.

Merge validation checked all 144 adapted matrices against their pre-merge values and all exported tensors for structure and finite values. BF16 merging can cause numerical differences from dynamically applying adapters. See merge_provenance.json and SHA256SUMS for export identities and validation.

Loading

from transformers import AutoModelForCausalLM, AutoTokenizer
import torch

model_id = "tanny2109/consensuslab-fact-checking-3b"
tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(model_id, torch_dtype=torch.bfloat16,
                                          device_map="auto")

Use tokenizer.apply_chat_template with the exact system/user messages from the study repository. Preserve the randomized mapping from A/B to semantic choices. The upstream Qwen Research license is included unchanged in LICENSE.

Downloads last month
-
Safetensors
Model size
3B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for tanny2109/consensuslab-fact-checking-3b

Collection including tanny2109/consensuslab-fact-checking-3b