Feature Extraction
Model2Vec
Safetensors
code
embeddings
static-embeddings
code-retrieval
dual-encoder
coir
hybrid
Eval Results (legacy)

miru-codev3-dual

Owned dual static code embedder for Miru / CoIR-style retrieval.

Two 256-d Model2Vec bags are fused at inference into a 512-d unit vector by score-fuse concat (α=0.5). Cosine on the concat equals the mean of the two branch cosines — no potion at inference.

Branch Hub / subdir Dim
A — Tokenlearn Zipf-SIF tokenlearn/ · miru-codev3-tokenlearn 256
B — distill_fuse distill_fuse/ · miru-codev3-distill-fuse 256
Dual concat this repo 512

Evaluation

CoIR / MTEB NDCG@10 (×100), self-reported:

  1. Dense — dual concat score-fuse (α=0.5), 10-task card average 40.78 (vs potion-code-16M-v2 39.08).
  2. Hybrid — dual ⊕ identifier BM25 with CoIR linear fusion dense_w=0.3, 10-task average 49.55 (vs potion+BM25 43.36).

The Hub eval widget lists both setups (task name includes dense vs hybrid).

Install

pip install model2vec numpy huggingface_hub

Usage

from dual_encode import DualConcatEncoder

enc = DualConcatEncoder.from_pretrained("Takara-DS1/miru-codev3-dual")
vecs = enc.encode(["def add(a, b): return a + b"])
assert vecs.shape[1] == 512

Fuse formula

e = normalize( concat( √α · e_A , √(1-α) · e_B ) )   # α=0.5
cos(e_q, e_d) = α · cos(A_q, A_d) + (1-α) · cos(B_q, B_d)

Hybrid (product / CoIR card)

score = 0.3 · minmax(dense) + 0.7 · minmax(bm25)

Recreate

See reproduce/REPRODUCE.md and reproduce/recreate.sh.

License

MIT

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Datasets used to train Takara-DS1/miru-codev3-dual

Evaluation results