Image-Text-to-Text
Transformers
Safetensors
English
gemma3n
function-calling
tool-use
on-device
mobile
gemma
litertlm
conversational
Instructions to use kontextdev/agent-gemma with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use kontextdev/agent-gemma with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("image-text-to-text", model="kontextdev/agent-gemma") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] pipe(text=messages)# Load model directly from transformers import AutoProcessor, AutoModelForMultimodalLM processor = AutoProcessor.from_pretrained("kontextdev/agent-gemma") model = AutoModelForMultimodalLM.from_pretrained("kontextdev/agent-gemma", device_map="auto") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] inputs = processor.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=40) print(processor.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use kontextdev/agent-gemma with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "kontextdev/agent-gemma" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "kontextdev/agent-gemma", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker
docker model run hf.co/kontextdev/agent-gemma
- SGLang
How to use kontextdev/agent-gemma with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "kontextdev/agent-gemma" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "kontextdev/agent-gemma", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "kontextdev/agent-gemma" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "kontextdev/agent-gemma", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }' - Docker Model Runner
How to use kontextdev/agent-gemma with Docker Model Runner:
docker model run hf.co/kontextdev/agent-gemma
File size: 4,274 Bytes
d7d68cb 4f6ad08 d7d68cb da7f35e 4f6ad08 d7d68cb 4f6ad08 d7d68cb 4f6ad08 da7f35e 4f6ad08 da7f35e 4f6ad08 da7f35e 4f6ad08 da7f35e 4f6ad08 d7d68cb 4f6ad08 da7f35e d7d68cb 4f6ad08 d7d68cb 4f6ad08 d7d68cb 4f6ad08 d7d68cb 4f6ad08 d7d68cb da7f35e 4f6ad08 d7d68cb 4f6ad08 d7d68cb 4f6ad08 da7f35e 4f6ad08 da7f35e 4f6ad08 da7f35e 4f6ad08 da7f35e 4f6ad08 da7f35e 4f6ad08 da7f35e 4f6ad08 da7f35e 4f6ad08 da7f35e d7d68cb 4f6ad08 d7d68cb 4f6ad08 d7d68cb 4f6ad08 d7d68cb 4f6ad08 d7d68cb 4f6ad08 d7d68cb 4f6ad08 d7d68cb 4f6ad08 d7d68cb 4f6ad08 d7d68cb 4f6ad08 d7d68cb 4f6ad08 d7d68cb 4f6ad08 d7d68cb 4f6ad08 d7d68cb 4f6ad08 d7d68cb da7f35e | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 | ---
language:
- en
license: gemma
library_name: transformers
base_model: google/gemma-3n-E2B-it
tags:
- function-calling
- tool-use
- on-device
- mobile
- gemma
- litertlm
---
# Agent Gemma β Gemma 3n E2B Fine-Tuned for Function Calling
A fine-tuned version of [google/gemma-3n-E2B-it](https://huggingface.co/google/gemma-3n-E2B-it) trained for on-device function calling using Google's [FunctionGemma](https://ai.google.dev/gemma/docs/functiongemma/function-calling-with-hf) technique.
## What's Different from Stock Gemma 3n
### Fixed: `format_function_declaration` Template Error
The stock Gemma 3n chat template uses `format_function_declaration()` β a custom Jinja2 function available in Google's Python tokenizer but **not supported by LiteRT-LM's on-device template engine**. This causes:
```
Failed to apply template: unknown function: format_function_declaration is unknown (in template:21)
```
This model replaces the stock template with a **LiteRT-LM compatible** template that uses only standard Jinja2 features (`tojson` filter, `<start_function_declaration>` / `<end_function_declaration>` markers). The template is embedded in both `tokenizer_config.json` and `chat_template.jinja`.
### Function Calling Format
The model uses the FunctionGemma markup format:
```
<start_function_call>call:function_name{param:<escape>value<escape>}<end_function_call>
```
Tool declarations are formatted as:
```
<start_function_declaration>{"name": "get_weather", "parameters": {...}}<end_function_declaration>
```
## Training Details
- **Base model:** google/gemma-3n-E2B-it (5.4B parameters)
- **Method:** QLoRA (rank=16, alpha=32) β 22.9M trainable parameters (0.42%)
- **Dataset:** [google/mobile-actions](https://huggingface.co/datasets/google/mobile-actions) (8,693 training samples)
- **Training:** 500 steps, batch_size=1, max_seq_length=512, learning_rate=2e-4
- **Precision:** bfloat16
## Usage
### With LiteRT-LM on Android (Kotlin)
```kotlin
// After converting to .litertlm format
val engine = Engine(EngineConfig(modelPath = "agent-gemma.litertlm"))
engine.initialize()
val conversation = engine.createConversation(
ConversationConfig(
systemMessage = Message.of("You are a helpful assistant."),
tools = listOf(MyToolSet()) // @Tool annotated class
)
)
// No format_function_declaration error!
conversation.sendMessageAsync(Message.of("What's the weather?"))
.collect { print(it) }
```
### With Transformers (Python)
```python
from transformers import AutoModelForCausalLM, AutoTokenizer
model = AutoModelForCausalLM.from_pretrained("kontextdev/agent-gemma")
tokenizer = AutoTokenizer.from_pretrained("kontextdev/agent-gemma")
messages = [
{"role": "developer", "content": "You are a helpful assistant."},
{"role": "user", "content": "What's the weather in Tokyo?"}
]
tools = [{"function": {"name": "get_weather", "parameters": {"type": "object", "properties": {"location": {"type": "string"}}}}}]
text = tokenizer.apply_chat_template(messages, tools=tools, tokenize=False, add_generation_prompt=True)
inputs = tokenizer(text, return_tensors="pt")
output = model.generate(**inputs, max_new_tokens=256)
print(tokenizer.decode(output[0]))
```
## Chat Template
The custom chat template (in `tokenizer_config.json` and `chat_template.jinja`) supports these roles:
- `developer` / `system` β system instructions + tool declarations
- `user` β user messages
- `model` / `assistant` β model responses, including `tool_calls`
- `tool` β tool execution results
## Converting to .litertlm
Use the [LiteRT-LM](https://github.com/google-ai-edge/LiteRT-LM) conversion tools to package for on-device deployment:
```bash
# The chat_template.jinja is included in this repo
python scripts/convert-to-litertlm.py \
--model_dir kontextdev/agent-gemma \
--output agent-gemma.litertlm
```
## Files
- `model-*.safetensors` β Merged model weights (bfloat16)
- `tokenizer_config.json` β Tokenizer config with embedded chat template
- `chat_template.jinja` β Standalone chat template file
- `config.json` β Model architecture config
- `checkpoint-*` β Training checkpoints (LoRA)
## License
This model inherits the [Gemma license](https://ai.google.dev/gemma/terms) from the base model.
|