Sakura-EmbeddingGemma-2-AutoRound / quantization_config.json
webmp3's picture
v2: re-quantized W4A16 (group 64, 200 iters, 234 real retrieval calibration texts); modest gains in BF16 fidelity and BEIR; README re-measured, build manifest added. v1 stays in git history.
3f8b39e verified
Raw History Blame Contribute Delete
290 Bytes
{
"bits": 4,
"data_type": "int",
"group_size": 64,
"sym": true,
"low_gpu_mem_usage": true,
"autoround_version": "0.16.0",
"block_name_to_quantize": "language_model.layers",
"quant_method": "auto-round",
"packing_format": "auto_round:auto_gptq",
"iters": 200
}