I needed a model that does reliable matching between English and the 3 Japanese writting system, so I calibrated the compressed one with English/Japanese datasets as well as my private dataset. I included a script in extras, private dataset is not included but generation code is there for reference. Attention layers are left uncompressed for precision and this makes a noticable difference.

Downloads last month
208
Safetensors
Model size
3B params
Tensor type
F32
BF16
F8_E4M3
U8
Inference Providers NEW
This model isn't deployed by any Inference Provider. 馃檵 Ask for provider support

Model tree for catplusplus/Qwen3-Embedding-4B-NVFP4

Quantized
(47)
this model