Instructions to use Lythri/Lythri-4B-A2B with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use Lythri/Lythri-4B-A2B with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="Lythri/Lythri-4B-A2B") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] pipe(text=messages)# Load model directly from transformers import AutoProcessor, AutoModelForMultimodalLM processor = AutoProcessor.from_pretrained("Lythri/Lythri-4B-A2B") model = AutoModelForMultimodalLM.from_pretrained("Lythri/Lythri-4B-A2B", device_map="auto") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] inputs = processor.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=40) print(processor.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use Lythri/Lythri-4B-A2B with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "Lythri/Lythri-4B-A2B" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Lythri/Lythri-4B-A2B", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/Lythri/Lythri-4B-A2B
- SGLang
How to use Lythri/Lythri-4B-A2B with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "Lythri/Lythri-4B-A2B" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Lythri/Lythri-4B-A2B", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "Lythri/Lythri-4B-A2B" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Lythri/Lythri-4B-A2B", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Docker Model Runner
How to use Lythri/Lythri-4B-A2B with Docker Model Runner:
docker model run hf.co/Lythri/Lythri-4B-A2B
Hugging Face | ModelScope | GitHub | Technical Report (coming soon)
Lythri
Introduction
Lythri is a family of on-device language models built for emotional companionship. Instead of chasing math and coding scores, Lythri is trained to understand how people feel and to hold natural, multi-turn conversations, while staying small enough to run locally on a laptop or phone.
| Model | Total Params | Active Params | Base Model | GGUF |
|---|---|---|---|---|
| Lythri-7B-A4B | 7.46B | 4.5B | Gemma 4 E4B | Lythri-7B-A4B-GGUF |
| Lythri-4B-A2B | 4.63B | 2.3B | Gemma 4 E2B | Lythri-4B-A2B-GGUF |
Emotional Intelligence (Preliminary)
Zero-shot results.
| Model | GoEmotions (Macro F1) | EmoBench (Acc) | EQ-Bench (v2) |
|---|---|---|---|
| Gemma 4 E4B | 8.00 | 31.39 | 43.19 |
| Lythri-7B-A4B | 31.54 | 46.83 | 49.41 |
Note: These are preliminary results from an internal evaluation that is not yet formal or complete, and they may differ slightly from the final numbers. For full details and authoritative results, including Lythri-4B-A2B, please refer to the technical report (coming soon).
General Benchmarks
| Benchmark | Lythri-4B-A2B | Lythri-7B-A4B |
|---|---|---|
| Knowledge | ||
| MMLU | 55.23 | 69.09 |
| MMLU-Pro | 24.42 | 38.17 |
| ARC-E | 81.40 | 83.42 |
| ARC-C | 52.99 | 60.58 |
| Reasoning | ||
| PIQA | 79.49 | 81.88 |
| HellaSwag | 72.80 | 78.29 |
| WinoGrande | 68.43 | 74.90 |
| General | ||
| CommonsenseQA | 65.52 | 77.07 |
| SocialIQA | 49.80 | 50.46 |
| TruthfulQA MC2 | 46.16 | 49.66 |
| Science | ||
| OpenBookQA | 41.00 | 43.60 |
| GPQA Diamond | 28.79 | 27.78 |
| Math | ||
| GSM8K | 28.81 | 62.02 |
| MATH | 3.62 | 21.28 |
| Reading | ||
| BoolQ | 73.15 | 85.32 |
| Code / Instruction | ||
| HumanEval | 28.66 | 45.12 |
| IFEval | 26.43 | 31.05 |
All benchmarks are evaluated with their official standard settings and in generative mode with chat template applied, reflecting real-world inference conditions. Think-tag outputs from model are stripped before answer extraction.
Quickstart
from transformers import AutoModelForCausalLM, AutoTokenizer
import torch
model_path = "Lythri/Lythri-7B-A4B" # or "Lythri/Lythri-4B-A2B"
tok = AutoTokenizer.from_pretrained(model_path)
model = AutoModelForCausalLM.from_pretrained(
model_path, dtype=torch.bfloat16, device_map="auto"
)
messages = [{"role": "user", "content": "My friend just lost their job and seems really down. What should I say to them?"}]
chat = tok.apply_chat_template(messages, tokenize=False, add_generation_prompt=True) + "<think>"
inputs = tok(chat, return_tensors="pt").to(model.device)
with torch.inference_mode():
out = model.generate(
**inputs,
max_new_tokens=2048,
do_sample=False,
eos_token_id=[1, 106],
)
print(tok.decode(out[0][inputs.input_ids.shape[1]:], skip_special_tokens=True))
Recommended Generation Config
generation_config = {
"temperature": 0.95,
"top_p": 0.9,
"top_k": 64,
"max_new_tokens": 2048,
"repetition_penalty": 1.05,
"do_sample": True,
"eos_token_id": [1, 106],
}
out = model.generate(**inputs, **generation_config)
Compute
The full development of Lythri, including training and evaluation, used about 2,842 GPU hours on NVIDIA RTX 6000D GPUs.
Limitations
- Lythri is optimized for conversation and emotional understanding, not for math, coding or complex reasoning.
- Lythri is not a substitute for professional mental health support. If you or someone you know is in crisis, please contact local emergency services or a crisis helpline.
- Like all language models, it can produce inaccurate or inappropriate content.
License
Lythri is built on Gemma 4 and is released under the Apache License 2.0.
- Downloads last month
- 868
Model tree for Lythri/Lythri-4B-A2B
Collection including Lythri/Lythri-4B-A2B
Evaluation results
- openai/gsm8k · Gsm8k View evaluation results leaderboard 28.81 *
- Idavidrein/gpqa · Diamond View evaluation results leaderboard 28.79 *
- TIGER-Lab/MMLU-Pro · Mmlu Pro View evaluation results leaderboard 24.42 *