Instructions to use utahnlp/laconic-8b with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use utahnlp/laconic-8b with Transformers:
# pip install -U transformers accelerate # Load model directly from transformers import AutoTokenizer, LlamaBiForMNTP tokenizer = AutoTokenizer.from_pretrained("utahnlp/laconic-8b") model = LlamaBiForMNTP.from_pretrained("utahnlp/laconic-8b", device_map="auto") - Notebooks
- Google Colab
- Kaggle
Integrate with Sentence Transformers via SparseEncoder
#1
by tomaarsen HF Staff - opened
Hello @brutusxu and team!
Also see https://huggingface.co/utahnlp/laconic-1b/discussions/1.
Pull Request overview
- Add Sentence Transformers support through
SparseEncoder - Include tokenizer files and a model card with usage examples
Details
I've added the configuration, tokenizer, and a minimal LlamaForMaskedLM shim to load LACONIC with bidirectional attention and SPLADE pooling. The largest observed difference from the original implementation was approximately 2.98e-6 for the example's scores in float32.
You can try this PR with sentence-transformers>=5.4.0 and transformers>=5.2.0:
from sentence_transformers import SparseEncoder
model = SparseEncoder("utahnlp/laconic-8b", trust_remote_code=True, revision="refs/pr/1")
queries = ["What is the capital of France?", "How do plants make food?"]
documents = [
"Paris is the capital and largest city of France.",
"Plants use sunlight to turn carbon dioxide and water into sugars through photosynthesis.",
"The piano is a musical instrument with a keyboard.",
]
query_embeddings = model.encode_query(queries, max_active_dims=512)
document_embeddings = model.encode_document(documents, max_active_dims=512)
print(query_embeddings.shape, document_embeddings.shape)
# torch.Size([2, 128256]) torch.Size([3, 128256])
print(model.sparsity(query_embeddings))
# {'active_dims': 48.0, 'sparsity_ratio': 0.999625748502994}
print(model.sparsity(document_embeddings))
# {'active_dims': 279.0, 'sparsity_ratio': 0.9978246631736527}
scores = model.similarity(query_embeddings, document_embeddings)
print(scores.cpu())
# tensor([[16.3916, 0.3913, 1.1160],
# [ 0.5188, 15.8925, 0.9306]])
My changes here are largely additive, so the existing usage via your own GitHub is not affected. I also updated the Model Card to make it clearer what this model is and does.
- Tom Aarsen
tomaarsen changed pull request status to open