Hugging Face's logo Hugging Face
  • Models
  • Datasets
  • Spaces
  • Buckets new
  • Docs
  • Enterprise
  • Pricing
    • Website
      • Tasks
      • HuggingChat
      • Collections
      • Languages
      • Organizations
    • Community
      • Blog
      • Posts
      • Daily Papers
      • Hardware
      • Learn
      • Discord
      • Forum
      • GitHub
    • Solutions
      • Team & Enterprise
      • Hugging Face PRO
      • Enterprise Support
      • Inference Providers
      • Inference Endpoints
      • Storage Buckets

  • Log In
  • Sign Up

webmp3
/
Sakura-EmbeddingGemma-2-AutoRound

Feature Extraction
sentence-transformers
Safetensors
Transformers
embedding_gemma2
sentence-similarity
autoround
4-bit precision
quantization
embeddinggemma
embedding
mrl
matryoshka
auto-round
Model card Files Files and versions
xet
Community

Instructions to use webmp3/Sakura-EmbeddingGemma-2-AutoRound with libraries, inference providers, notebooks, and local apps. Follow these links to get started.

  • Libraries
  • sentence-transformers

    How to use webmp3/Sakura-EmbeddingGemma-2-AutoRound with sentence-transformers:

    from sentence_transformers import SentenceTransformer
    
    model = SentenceTransformer("webmp3/Sakura-EmbeddingGemma-2-AutoRound")
    
    sentences = [
        "The weather is lovely today.",
        "It's so sunny outside!",
        "He drove to the stadium."
    ]
    embeddings = model.encode(sentences)
    
    similarities = model.similarity(embeddings, embeddings)
    print(similarities.shape)
    # [3, 3]
  • Transformers

    How to use webmp3/Sakura-EmbeddingGemma-2-AutoRound with Transformers:

    # Use a pipeline as a high-level helper
    from transformers import pipeline
    
    pipe = pipeline("feature-extraction", model="webmp3/Sakura-EmbeddingGemma-2-AutoRound")
    # pip install -U transformers accelerate
    # Load model directly
    from transformers import AutoProcessor, AutoModel
    
    processor = AutoProcessor.from_pretrained("webmp3/Sakura-EmbeddingGemma-2-AutoRound")
    model = AutoModel.from_pretrained("webmp3/Sakura-EmbeddingGemma-2-AutoRound", device_map="auto")
  • Notebooks
  • Google Colab
  • Kaggle
Sakura-EmbeddingGemma-2-AutoRound / build
70.7 kB
Ctrl+K
Ctrl+K
  • 1 contributor
History: 1 commit
webmp3's picture
webmp3
v2: re-quantized W4A16 (group 64, 200 iters, 234 real retrieval calibration texts); modest gains in BF16 fidelity and BEIR; README re-measured, build manifest added. v1 stays in git history.
3f8b39e verified 2 days ago
  • calibration_dual256.json
    63.4 kB
    v2: re-quantized W4A16 (group 64, 200 iters, 234 real retrieval calibration texts); modest gains in BF16 fidelity and BEIR; README re-measured, build manifest added. v1 stays in git history. 2 days ago
  • requant.py
    7.37 kB
    v2: re-quantized W4A16 (group 64, 200 iters, 234 real retrieval calibration texts); modest gains in BF16 fidelity and BEIR; README re-measured, build manifest added. v1 stays in git history. 2 days ago