ML-LaySum Gemma 2 2B Instruct SFT

Full-parameter supervised fine-tuning of Gemma 2 2B Instruct on ML-LaySum for English lay summarization of machine-learning research. The input contains the paper's abstract, introduction, and conclusion where available; the target is its reference lay summary.

Training

Training uses 6,911 examples and 879 validation examples. The dataset has 852 held-out test examples. The final checkpoint is published after 1,296 optimizer steps (3 epochs); it is selected by completion, not best validation loss. Prompt tokens are masked from the supervised loss. Missing paper sections are not synthesized.

Training uses float32 master parameters with bfloat16 autocast. The published weights use bfloat16 (5.23 GB of weight files, excluding tokenizer and metadata). Export conversion does not change the training model's parameter precision. Optimizer state is not included.

The AdamW optimizer distributes its states using PyTorch ZeroRedundancyOptimizer across DDP ranks. All model parameters remain trainable, and the export contains the complete model weights.

Setting Value
epochs 3
learning rate 2e-05
micro batch size 1
gradient accumulation steps 2
max length 8192
seed 42
warmup ratio 0.03
weight decay 0.0
bf16 True
optimizer AdamW
optimizer state sharding PyTorch ZeroRedundancyOptimizer across DDP ranks
world size 8
effective batch size 16

Long-input policy: trim_introduction. Only the suffix of an overlong introduction is removed from the model input, preserving the complete abstract, closing section when present, and reference summary. All records are retained, and the released dataset is unchanged. This affected 19 train examples, 2 validation examples.

The sanitized training_summary.json records settings, split-count provenance, validation losses, and SHA-256 digests of the inference artifacts.

Usage

Load with a Transformers release supporting Gemma 2. Apply the included chat template with a user message that asks for a lay summary and supplies the paper sections. Generate only the continuation after the prompt. Training uses all model parameters; this repository contains a complete model, not a LoRA adapter, and can initialize subsequent reinforcement-learning training.

Evaluation and limitations

This publication script does not evaluate the held-out test split. Training and validation losses do not establish factual accuracy or accessibility for non-experts. Generated summaries may omit qualifications or introduce unsupported statements; evaluation should check both source faithfulness and reader comprehension. No state-of-the-art claim is made.

License and artifacts

Use and redistribution are governed by the Gemma Terms of Use. Dataset and source-paper terms apply separately; see ML-LaySum. Only inference weights, tokenizer/configuration files, and sanitized metadata are distributed. Optimizer and random-state checkpoints are not included; resuming the original optimizer state is therefore not supported by this export.

Downloads last month
117
Safetensors
Model size
3B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for sandeep123/ML-LaySum-Gemma2-2B-Instruct-SFT

Finetuned
(1109)
this model