File size: 1,612 Bytes
61146e3
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
20048fc
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
---
license: mit
base_model: google/gemma-2b
tags:
- text-generation-inference
- transformers
- gemma
- mini-gemma
- agentic-ai
model_type: gemma
pipeline_tag: text-generation
---

# Mini-Gemma Custom Model

This repository contains a custom domain-specialized fine-tune of the Gemma architecture, optimized for specific text distributions and patterns. The model was trained using the Hugging Face `Trainer` on an accelerated NVIDIA GPU cluster.

## ๐Ÿ“Š Training Performance & Metrics

The model successfully converged over its training run with highly stable gradients:
* **Total Training Steps:** 20,000
* **Final Total Train Loss:** `3.478`
* **Final Step Loss:** `2.988`
* **Gradient Norm Stability:** Stable at `~1.12`
* **Training Status:** Complete / Fully Converged

## ๐Ÿš€ Quick Start & Usage

You can easily load and run this model locally using the Transformers library:

```python
import torch
from transformers import AutoTokenizer, AutoModelForCausalLM, pipeline

model_id = "agentbyumer/mini-gemma"

tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(
    model_id,
    torch_dtype=torch.float16,
    device_map="auto"
)

generator = pipeline("text-generation", model=model, tokenizer=tokenizer)

prompt = "Your specialized prompt here"
outputs = generator(
    prompt, 
    max_new_tokens=150, 
    do_sample=True, 
    temperature=0.7,
    return_full_text=False
)
print(outputs[0]['generated_text'])
```

## ๐Ÿ“œ License

This project is licensed under the permissive MIT License. See the accompanying [LICENSE](./LICENSE) file for full details.