Instructions to use Prashanth-24/gemma-4-e4b-python-lora with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- MLX
How to use Prashanth-24/gemma-4-e4b-python-lora with MLX:
# Make sure mlx-lm is installed # pip install --upgrade mlx-lm # if on a CUDA device, also pip install mlx[cuda] # Generate text with mlx-lm from mlx_lm import load, generate model, tokenizer = load("Prashanth-24/gemma-4-e4b-python-lora") prompt = "Once upon a time in" text = generate(model, tokenizer, prompt=prompt, verbose=True) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- LM Studio
- MLX LM
How to use Prashanth-24/gemma-4-e4b-python-lora with MLX LM:
Generate or start a chat session
# Install MLX LM uv tool install mlx-lm # Generate some text mlx_lm.generate --model "Prashanth-24/gemma-4-e4b-python-lora" --prompt "Once upon a time"
gemma-4-e4b-python-lora
A LoRA adapter that fine-tunes mlx-community/gemma-4-e4b-4bit
for Python code generation, trained with mlx_lm.
This repo contains only the adapter weights — you load it on top of the base model, you don't need to download a separate fused model.
Training details
| Base model | mlx-community/gemma-4-e4b-4bit |
| Dataset | iamtarun/python_code_instructions_18k_alpaca |
| Fine-tune type | LoRA |
| LoRA rank | 8 |
| LoRA scale | 20 |
| LoRA dropout | 0.0 |
| Layers adapted | 16 |
| Learning rate | 1e-5 |
| Optimizer | Adam |
| Iterations | 500 |
Training examples were formatted as:
### User:
{instruction}
### Assistant:
{output}
Usage
mlx_lm.load needs the adapter on disk, so download it locally first, then load it
exactly as you would a local adapter:
pip install -U mlx-lm huggingface_hub
from huggingface_hub import snapshot_download
from mlx_lm import load, generate
adapter_path = snapshot_download("Prashanth-24/gemma-4-e4b-python-lora")
model, tokenizer = load("mlx-community/gemma-4-e4b-4bit", adapter_path=adapter_path)
prompt = "### User:\nwrite a python code to ADD two numbers\n\n### Assistant:\n"
print(generate(model, tokenizer, prompt=prompt, max_tokens=200))
Example output
Prompt: Create a python program to multiply two numbers.
Base model
Here's a response to your query:
**1. Introduction**
**2. Part I (Body)**
**Introduction**
The concept of unit multiplication is relevant in understanding how to multiply two numbers.
**Part II**
**Part II**
**Key Topics**
**Key Topics**
- **Factors to Consider**
- **Key Topics**
- **Key Topics**
...
Fine-tuned model
def multiply_two_numbers(a, b):
return a * b
print(multiply_two_numbers(2, 3))
The base model drifts into a generic essay outline and never writes any code. The
fine-tuned model answers directly with a correct, runnable Python function — the LoRA
clearly taught the ### Assistant: turn to produce code instead of chatter.
Limitations
This is a small run (500 iterations, rank 8). It reliably improves "answer with code" behavior and formatting, but it's a lightweight fine-tune — treat outputs as a draft to review, not production-ready code.
Quantized
Model tree for Prashanth-24/gemma-4-e4b-python-lora
Base model
mlx-community/gemma-4-e4b-4bit