alphabrothers/alpha-sts-v0

Korean sentence embedding model (568M parameters) specialised for semantic textual similarity (STS), fine-tuned from dragonkue/snowflake-arctic-embed-l-v2.0-ko.

Evaluation — MTEB(kor, v2), STS

Task Split Score (cosine Spearman)
KLUE-STS validation 91.46
KorSTS test 84.87
STS17 test 86.39
STS mean 87.57

Training data

Only the official train splits of the two Korean STS datasets below were used:

  • KLUE-STS (klue/klue, sts) train
  • KorSTS (dkoterwa/kor-sts) train

No evaluation split was used for training. Hyper-parameters were selected on a held-out development set (500 pairs held out from KLUE-STS train + KorSTS valid), not on the benchmark evaluation splits. These datasets are declared as training_datasets (KLUE-STS, KorSTS) in the MTEB model metadata, so the model is not zero-shot on these tasks. STS17 (ko-ko) has no train split and was not used.

Training details

  • Base: dragonkue/snowflake-arctic-embed-l-v2.0-ko
  • CoSENTLoss on scored pairs (score/5) · batch 64 · lr 2e-5 · cosine schedule, warmup 10% · 4 epochs · max_len 128 · CLS pooling · L2-normalised · bf16

Usage

from sentence_transformers import SentenceTransformer
model = SentenceTransformer("alphabrothers/alpha-sts-v0")
emb = model.encode(["비행기가 이륙하고 있다.", "항공기가 이륙 중이다."])
print(model.similarity(emb[0], emb[1]))

No prompt / instruction is needed.

Limitations

The model is optimised for sentence-level similarity. Retrieval and classification quality may be lower than the base model.

License

apache-2.0 (same as the base model dragonkue/snowflake-arctic-embed-l-v2.0-ko).

Downloads last month
10
Safetensors
Model size
0.6B params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for alphabrothers/alpha-sts-v0

Space using alphabrothers/alpha-sts-v0 1