Add Sentence Transformers usage

#16
by tomaarsen HF Staff - opened
Files changed (1) hide show
  1. README.md +36 -0
README.md CHANGED
@@ -4,6 +4,8 @@ language:
4
  - en
5
  tags:
6
  - ColBERT
 
 
7
  ---
8
  <p align="center">
9
  <img align="center" src="https://github.com/stanford-futuredata/ColBERT/blob/main/docs/images/colbertofficial.png?raw=true" width="430px" />
@@ -43,6 +45,40 @@ These rich interactions allow ColBERT to surpass the quality of _single-vector_
43
 
44
  ----
45
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
46
  ## ColBERTv1
47
 
48
  The ColBERTv1 code from the SIGIR'20 paper is in the [`colbertv1` branch](https://github.com/stanford-futuredata/ColBERT/tree/colbertv1). See [here](#branches) for more information on other branches.
 
4
  - en
5
  tags:
6
  - ColBERT
7
+ - multi-vector
8
+ - sentence-transformers
9
  ---
10
  <p align="center">
11
  <img align="center" src="https://github.com/stanford-futuredata/ColBERT/blob/main/docs/images/colbertofficial.png?raw=true" width="430px" />
 
45
 
46
  ----
47
 
48
+ ## Usage with Sentence Transformers
49
+
50
+ This model can be used with [Sentence Transformers](https://www.sbert.net/) as a multi-vector (ColBERT-style late interaction) retriever via the `MultiVectorEncoder`:
51
+
52
+ ```bash
53
+ pip install "sentence-transformers>=6.0.0"
54
+ ```
55
+
56
+ ```python
57
+ from sentence_transformers import MultiVectorEncoder
58
+
59
+ model = MultiVectorEncoder("colbert-ir/colbertv2.0")
60
+
61
+ query = "Which planet is known as the Red Planet?"
62
+ documents = [
63
+ "Venus is often called Earth's twin because of its similar size and proximity.",
64
+ "Mars, known for its reddish appearance, is often referred to as the Red Planet.",
65
+ "Jupiter, the largest planet in our solar system, has a prominent red spot.",
66
+ "Saturn, famous for its rings, is sometimes mistaken for the Red Planet.",
67
+ ]
68
+
69
+ query_embeddings = model.encode_query(query)
70
+ document_embeddings = model.encode_document(documents)
71
+ print(query_embeddings.shape, document_embeddings[0].shape)
72
+ # (32, 128) (17, 128)
73
+
74
+ # MaxSim late-interaction scoring (higher is more relevant)
75
+ scores = model.similarity(query_embeddings, document_embeddings)
76
+ print(scores)
77
+ # tensor([[12.7970, 27.1945, 23.8495, 24.5656]])
78
+ ```
79
+
80
+ ----
81
+
82
  ## ColBERTv1
83
 
84
  The ColBERTv1 code from the SIGIR'20 paper is in the [`colbertv1` branch](https://github.com/stanford-futuredata/ColBERT/tree/colbertv1). See [here](#branches) for more information on other branches.