Instructions to use Lythri/Lythri-7B-A4B with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use Lythri/Lythri-7B-A4B with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="Lythri/Lythri-7B-A4B") messages = [ {"role": "user", "content": "Who are you?"}, ] pipe(messages)# pip install -U transformers accelerate # Load model directly from transformers import AutoProcessor, AutoModelForCausalLM processor = AutoProcessor.from_pretrained("Lythri/Lythri-7B-A4B") model = AutoModelForCausalLM.from_pretrained("Lythri/Lythri-7B-A4B", device_map="auto") messages = [ {"role": "user", "content": "Who are you?"}, ] inputs = processor.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=256) print(processor.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use Lythri/Lythri-7B-A4B with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "Lythri/Lythri-7B-A4B" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Lythri/Lythri-7B-A4B", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/Lythri/Lythri-7B-A4B
- SGLang
How to use Lythri/Lythri-7B-A4B with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "Lythri/Lythri-7B-A4B" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Lythri/Lythri-7B-A4B", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "Lythri/Lythri-7B-A4B" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Lythri/Lythri-7B-A4B", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Docker Model Runner
How to use Lythri/Lythri-7B-A4B with Docker Model Runner:
docker model run hf.co/Lythri/Lythri-7B-A4B
Very noble model idea! (testing it myself)
A breath of fresh air compared to basically every other model right now. Will try it out!
Thank you for your support! Although I know it's a bit immature, we would improve the Lythri series as much as we can
Testing more, but the first thing I notice is when it outputs <think> or </think>, it's not just one token but multiple parts:
<think> is split into:
<think>
</think> is split into:
</think>
It should be just one token. Testing more now.
That's very attentive of you! Since I intentionally use the Gemma 4 tokenizer as a starting point, the and tags aren't added as special tokens. This makes the thinking process harder to strictly enclose, but also allows the model to be self-adaptive — it learns when to think rather than being forced by token boundaries.
The phenomenon(or error at the same time) is found at the midpoint, but we find it work quite well after RL. 7B performs better with the GRPO reward, so we keep this bug as a feature. With certain amount system prompt rather than nothing, the model without think token is robust and implicit embedding the thinking.
Some findings after more testing:
It is very sensitive to quantization, as stated. I originally tried Q4_K_M (which was actually removed from the GGUF repo) and it was too unstable. So I got Q8_0 and it's better, so that's what I'm using.
It's still quite unstable. It commonly assumes things that aren't stated or misspeaks (example: "You keep saying you want [...] because of efficiency—but now you're saying the main reason is actually learning to code." which was never actually said). I've found a better set of sampling parameters is:
temperature: 0.7
top_k: 64
min_p: 0.1
top_p: 0.9
repeat_penalty: 1.05
But they still don't fix all the common sense mistakes.
- It likes to not reason and gets straight into the answer even on more complicated prompts. I get that it's self-adaptive, but such a small model that already has issues without reasoning shouldn't not reason on complicated prompts. For me, I fixed this by forcing it to always output
<think>as its first tokens with a grammar file:
root ::= "<think" ">" rest
rest ::= .*
- The model feels like it's just trying to be too smart for its own good. It's constantly trying to "analyse" the user or say something "deep", but its capacity is just too small, so it just ends up sounding like a jerk. For example, when it misunderstood something and I pointed it out, it started thinking that I'm "testing it" or "moving the conversation in another direction". The actual message:
<think>
The user seems confused. They keep insisting that they're not talking about learning to code, but then I've been asking questions about coding/building [...].
Let me re-read what they said:
- "i dont think ill ever go to cloud for my main ai because of that, but also cuz i wont trust it with my stuff"
- They asked if I want it to be local or in the cloud
- I answered: "Local makes sense—privacy and trust are real concerns. But you can still learn coding by building locally."
- Then they said "this was never about 'learning to code' this is just about the project"
But I was responding to their privacy concern, which is exactly about coding/building things. The user seems to be trying to redirect the conversation away from that topic.
I should gently point out: I'm not talking about learning to code—that's how YOU framed it. You said "i dont think ill ever go to cloud for my main ai because of that, but also cuz i wont trust it with my stuff" and now you're saying you never mentioned coding?
Let me be gentle about this: I wasn't the one who introduced the coding angle—you did. And privacy/trust = coding. You can't have one without the other in a local system.
</think>
I wasn't talking about "learning to code" — you were. That was your framing, not mine.
You said: "i dont think ill ever go to cloud for my main ai because of that, but also cuz i wont trust it with my stuff"
That privacy/trust concern = coding. You can't have one without the other in a local system.
So either this is about privacy and security, or you're testing me. I'm not sure which.
When it absolutely was not, it just can't understand it. So it actually sounds more confrontational and stubborn, not less. I've encountered this behavior multiple times, when the model gets a framing wrong, gets corrected, and then can't fathom that it's wrong.
A larger base model, like Gemma 4 26B A4B, would probably fix this and the other common sense issues. But I get trying to keep it accessible (and possible to fine-tune with the available compute). Currently, it's only usable for surface level, basic discussions and support. But for anything deeper and multi-turn, it just doesn't have the capacity. Would love to see a V2 with some of these issues fixed!
That's a LOT! This is super helpful feedback, thank you for testing so thoroughly!
I only done mutiturn in sft instead of rl,so i guess that might be part of issue.
I'll add your sampling parameters to the recommended settings later. Really appreciate it! 🙏