Sentence Similarity
Safetensors
Japanese
RAGatouille
bert
ColBERT

Add Sentence Transformers usage

#1
by tomaarsen HF Staff - opened

Hello!

As of Sentence Transformers v6.0.0, this checkpoint loads directly as a multi-vector (ColBERT-style late interaction) retriever through the new MultiVectorEncoder. This PR adds a Sentence Transformers usage section to the model card and the multi-vector and sentence-transformers tags. The weights and the existing usage are untouched.

pip install "sentence-transformers>=6.0.0" fugashi unidic-lite
from sentence_transformers import MultiVectorEncoder

model = MultiVectorEncoder("bclavie/JaColBERT")

query = "日本で一番高い山は何ですか?"
documents = [
    "富士山は日本で最も高い山で、標高は3776メートルです。",
    "東京は日本の首都で、世界最大の都市圏の一つです。",
]

query_embeddings = model.encode_query(query)
document_embeddings = model.encode_document(documents)
print(query_embeddings.shape, document_embeddings[0].shape)
# torch.Size([32, 128]) torch.Size([21, 128])

# MaxSim late-interaction scoring (higher is more relevant)
scores = model.similarity(query_embeddings, document_embeddings)
print(scores)
# tensor([[28.1927, 16.7800]], device='cuda:0')

No configuration files are needed. config.json already declares architectures: ["HF_ColBERT"], so Sentence Transformers takes its Stanford-NLP ColBERT path and reads the [unused0] / [unused1] markers, the 32-token query expansion, the 228-token document length, the punctuation masking and the 128-dim projection (the root linear.weight) straight out of the checkpoint and artifact.metadata.

Verified against a colbert-ai reference running the transformers version recorded in config.json (4.36.2): per-token cosine similarity 1.000000 on the query and on all four test documents, identical token counts, and MaxSim scores matching to 1.9e-06. Queries longer than 32 tokens and documents longer than 228 tokens truncate identically, and the punctuation skiplist drops the same tokens. Loading by repo id from the Hub reproduces the same numbers today, so the model card is the only thing this PR changes.

Happy to tweak anything you would like changed. Please let me know if you have any questions or feedback!

  • Tom Aarsen
tomaarsen changed pull request status to open
Ready to merge
This branch is ready to get merged automatically.

Sign up or log in to comment