File size: 3,833 Bytes
e3ecf59 5122aa8 e3ecf59 5122aa8 e3ecf59 5122aa8 e3ecf59 5122aa8 | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 | ---
title: README
emoji: π
colorFrom: indigo
colorTo: blue
sdk: static
pinned: false
license: mit
---
<!--
QuerynAi β Hugging Face organization card.
Local draft (git-ignored). Paste into the org card at
https://huggingface.co/organizations/QuerynAi/settings β or the QuerynAi/README repo.
Fill the <β¦> placeholders before publishing.
-->
# Queryn β Embedding Translation
**Move a corpus between embedding models without re-embedding it.** Given a text
chunk's embedding in Model A's space, a Queryn adapter returns the equivalent
vector in Model B's space β so a vector index built with one model can be served
against another after a lightweight transform instead of a full, expensive
backfill.
## What's in this org
The **[Queryn Embedding Adapters](https://huggingface.co/QuerynAi)** collection β
one small ONNX model per directed model pair (`queryn-adapter-<source>_to_<target>`).
Each repo contains:
- `model.onnx` β the adapter, opset 17, dynamic batch axis. Runs anywhere
`onnxruntime` runs; no PyTorch needed.
- `model.safetensors` β the same weights, for retraining / inspection.
- `config.json` β dimensions, the I/O contract, and provenance.
- a model card with per-pair metrics and training plots.
Every adapter L2-normalizes its input and output internally, so you feed raw
embeddings straight in and get unit vectors back.
## Using an adapter
```python
import numpy as np, onnxruntime as ort
from huggingface_hub import hf_hub_download
repo = "QuerynAi/queryn-adapter-ada-002_to_bge-m3"
sess = ort.InferenceSession(hf_hub_download(repo, "model.onnx"),
providers=["CPUExecutionProvider"])
src = np.random.rand(8, 1536).astype(np.float32) # your ada-002 embeddings
tgt = sess.run(["target_embedding"], {"source_embedding": src})[0]
# tgt: (8, 1024) unit vectors in bge-m3 space
```
## How the adapters are built
- One **linear** projection and one **1-hidden-layer GELU MLP** (with a
compressed latent below both dimensions) are trained per pair; whichever scores
higher on held-out cosine similarity is published, ties to linear. Each pair's
card reports both, so the linear-vs-nonlinear trade-off is visible.
- Trained on a unified **~349,674-row, five-domain corpus** (arXiv abstracts,
Australian case law, SQuAD passages, PubMed RCT abstracts, financial/markets
news), optimizing `1 β mean cosine similarity` with Adam and LR scheduling.
- Some embedding spaces are near-isomorphic and align well with a single matrix;
others need the nonlinear map and still lose accuracy. Check the card metrics
for your pair before relying on it β translation is an approximation, not a
substitute for re-embedding when fidelity is critical.
## Model coverage
| Model | Dim | Role in v1 |
|---|---|---|
| `ada-002` | 1536 | source only (deprecated as a target) |
| `te3-small` | 1536 | source + target |
| `qwen3-emb-8b` | 4096 | source + target |
| `bge-m3` | 1024 | source + target |
| `me5-large` | 1024 | source + target |
| `pplx-embed-1` | 1024 | source + target |
| `nemotron-1b-free` | 2048 | source + target |
| `fastembed-bge-small` | 384 | source + target |
49 directed pairs in the current (v1) generation.
## Links
- **Code & pipeline:** <GitHub repo URL>
- **Training corpus + paired embeddings:** <Kaggle dataset URL>
- **Papers behind the approach:** Vec2Vec ([arXiv:2306.12689](https://arxiv.org/abs/2306.12689)),
mini-vec2vec ([arXiv:2510.02348](https://arxiv.org/abs/2510.02348))
## License
Adapter models: **MIT**. The Queryn codebase: **Apache-2.0**. Training data keeps
each source corpus's own license β see the dataset's `LICENSE` manifest.
## Status
An independent research project on practical embedding-space alignment. Issues and
findings welcome via the repo; the adapters are provided as-is. |