Mercan 0.8B Preference Plus β Q4_K_M
This repository contains the HH-RLHF-TR continuation-DPO Mercan 0.8B deployment artifact.
Main artifact
Mercan-0.8B-Preference-Plus-Q4_K_M.mercanβ Mercan v1 / GGUF v3, Q4_K_MMercan-0.8B-Preference-Plus-Q4_K_M.mercan.sha256β SHA-256Mercan-0.8B-Preference-Plus-Q4_K_M.mercan.jsonβ conversion/provenance manifest
Training chain: continuation SFT -> DPO -> HH-RLHF-TR continuation DPO
Source checkpoint: step_00004958.pt
Base DPO checkpoint: step_00007895.pt
Preference source: MercanAI/hh-rlhf-tr
Model contract
- tokenizer: NDSRF004
- tokenizer SHA-256:
72412d981dac65a29d1767bc98821fc2bcffc2de53c534e7c719598515bfb600 - vocab rows: 32,002
- chat control IDs: 32,000 / 32,001
- canonical roles:
sistem,kullanici,asistan - hidden size: 1536
- layers: 24
- attention heads: 12
- KV heads: 4
- FFN width: 5632
- context/sliding window: 32,768
- RoPE base: 1,000,000
- MorphFFN layers: 18
A fixed 100-question TR-MMLU diagnostic sample produced 26/100, matching the pre-HH DPO checkpoint on the same frozen sample.
The structural control IDs are stored with Mercan runtime aliases <|im_start|> / <|im_end|> while canonical role headers remain Turkish.
CLI
Use the explicit artifact name:
mercan run MercanAI/Mercan-0.8B-Preference-Plus-Q4:Mercan-0.8B-Preference-Plus-Q4_K_M.mercan