Instructions to use sandeep123/ML-LaySum-Gemma2-2B-Instruct-SFT with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use sandeep123/ML-LaySum-Gemma2-2B-Instruct-SFT with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="sandeep123/ML-LaySum-Gemma2-2B-Instruct-SFT") messages = [ {"role": "user", "content": "Who are you?"}, ] pipe(messages)# pip install -U transformers accelerate # Load model directly from transformers import AutoTokenizer, AutoModelForCausalLM tokenizer = AutoTokenizer.from_pretrained("sandeep123/ML-LaySum-Gemma2-2B-Instruct-SFT") model = AutoModelForCausalLM.from_pretrained("sandeep123/ML-LaySum-Gemma2-2B-Instruct-SFT", device_map="auto") messages = [ {"role": "user", "content": "Who are you?"}, ] inputs = tokenizer.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=256) print(tokenizer.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use sandeep123/ML-LaySum-Gemma2-2B-Instruct-SFT with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "sandeep123/ML-LaySum-Gemma2-2B-Instruct-SFT" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "sandeep123/ML-LaySum-Gemma2-2B-Instruct-SFT", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/sandeep123/ML-LaySum-Gemma2-2B-Instruct-SFT
- SGLang
How to use sandeep123/ML-LaySum-Gemma2-2B-Instruct-SFT with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "sandeep123/ML-LaySum-Gemma2-2B-Instruct-SFT" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "sandeep123/ML-LaySum-Gemma2-2B-Instruct-SFT", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "sandeep123/ML-LaySum-Gemma2-2B-Instruct-SFT" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "sandeep123/ML-LaySum-Gemma2-2B-Instruct-SFT", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Docker Model Runner
How to use sandeep123/ML-LaySum-Gemma2-2B-Instruct-SFT with Docker Model Runner:
docker model run hf.co/sandeep123/ML-LaySum-Gemma2-2B-Instruct-SFT
ML-LaySum Gemma 2 2B Instruct SFT
Full-parameter supervised fine-tuning of Gemma 2 2B Instruct on ML-LaySum for English lay summarization of machine-learning research. The input contains the paper's abstract, introduction, and conclusion where available; the target is its reference lay summary.
Training
Training uses 6,911 examples and 879 validation examples. The dataset has 852 held-out test examples. The final checkpoint is published after 1,296 optimizer steps (3 epochs); it is selected by completion, not best validation loss. Prompt tokens are masked from the supervised loss. Missing paper sections are not synthesized.
Training uses float32 master parameters with bfloat16 autocast. The published weights use bfloat16 (5.23 GB of weight files, excluding tokenizer and metadata). Export conversion does not change the training model's parameter precision. Optimizer state is not included.
The AdamW optimizer distributes its states using PyTorch ZeroRedundancyOptimizer across DDP ranks. All model parameters remain trainable, and the export contains the complete model weights.
| Setting | Value |
|---|---|
| epochs | 3 |
| learning rate | 2e-05 |
| micro batch size | 1 |
| gradient accumulation steps | 2 |
| max length | 8192 |
| seed | 42 |
| warmup ratio | 0.03 |
| weight decay | 0.0 |
| bf16 | True |
| optimizer | AdamW |
| optimizer state sharding | PyTorch ZeroRedundancyOptimizer across DDP ranks |
| world size | 8 |
| effective batch size | 16 |
Long-input policy: trim_introduction. Only the suffix of an overlong introduction is removed from the model input, preserving the complete abstract, closing section when present, and reference summary. All records are retained, and the released dataset is unchanged. This affected 19 train examples, 2 validation examples.
The sanitized training_summary.json records settings, split-count provenance, validation losses, and SHA-256 digests of the inference artifacts.
Usage
Load with a Transformers release supporting Gemma 2. Apply the included chat template with a user message that asks for a lay summary and supplies the paper sections. Generate only the continuation after the prompt. Training uses all model parameters; this repository contains a complete model, not a LoRA adapter, and can initialize subsequent reinforcement-learning training.
Evaluation and limitations
This publication script does not evaluate the held-out test split. Training and validation losses do not establish factual accuracy or accessibility for non-experts. Generated summaries may omit qualifications or introduce unsupported statements; evaluation should check both source faithfulness and reader comprehension. No state-of-the-art claim is made.
License and artifacts
Use and redistribution are governed by the Gemma Terms of Use. Dataset and source-paper terms apply separately; see ML-LaySum. Only inference weights, tokenizer/configuration files, and sanitized metadata are distributed. Optimizer and random-state checkpoints are not included; resuming the original optimizer state is therefore not supported by this export.
- Downloads last month
- 117