minishlab/tokenlearn-cornstack-docs-coderankembed-v2
Viewer • Updated • 600k • 33 • 1
How to use Takara-DS1/miru-codev3-dual with Model2Vec:
from model2vec import StaticModel
model = StaticModel.from_pretrained("Takara-DS1/miru-codev3-dual")
embeddings = model.encode(["It's dangerous to go alone!", "It's a secret to everybody."])
print(embeddings.shape)Owned dual static code embedder for Miru / CoIR-style retrieval.
Two 256-d Model2Vec bags are fused at inference into a 512-d unit vector by score-fuse concat (α=0.5). Cosine on the concat equals the mean of the two branch cosines — no potion at inference.
| Branch | Hub / subdir | Dim |
|---|---|---|
| A — Tokenlearn Zipf-SIF | tokenlearn/ · miru-codev3-tokenlearn |
256 |
| B — distill_fuse | distill_fuse/ · miru-codev3-distill-fuse |
256 |
| Dual concat | this repo | 512 |
CoIR / MTEB NDCG@10 (×100), self-reported:
dense_w=0.3, 10-task average 49.55 (vs potion+BM25 43.36).The Hub eval widget lists both setups (task name includes dense vs hybrid).
pip install model2vec numpy huggingface_hub
from dual_encode import DualConcatEncoder
enc = DualConcatEncoder.from_pretrained("Takara-DS1/miru-codev3-dual")
vecs = enc.encode(["def add(a, b): return a + b"])
assert vecs.shape[1] == 512
e = normalize( concat( √α · e_A , √(1-α) · e_B ) ) # α=0.5
cos(e_q, e_d) = α · cos(A_q, A_d) + (1-α) · cos(B_q, B_d)
score = 0.3 · minmax(dense) + 0.7 · minmax(bm25)
See reproduce/REPRODUCE.md and reproduce/recreate.sh.
MIT