Instructions to use ikppramesh/irx-2 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- MLX
How to use ikppramesh/irx-2 with MLX:
# Make sure mlx-lm is installed # pip install --upgrade mlx-lm # Generate text with mlx-lm from mlx_lm import load, generate model, tokenizer = load("ikppramesh/irx-2") prompt = "Write a story about Einstein" messages = [{"role": "user", "content": prompt}] prompt = tokenizer.apply_chat_template( messages, add_generation_prompt=True ) text = generate(model, tokenizer, prompt=prompt, verbose=True) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- LM Studio
- Pi
How to use ikppramesh/irx-2 with Pi:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "ikppramesh/irx-2"
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "mlx-lm": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "ikppramesh/irx-2" } ] } } }Run Pi
# Start Pi in your project directory: pi
- MLX LM
How to use ikppramesh/irx-2 with MLX LM:
Generate or start a chat session
# Install MLX LM uv tool install mlx-lm # Interactive chat REPL mlx_lm.chat --model "ikppramesh/irx-2"
Run an OpenAI-compatible server
# Install MLX LM uv tool install mlx-lm # Start the server mlx_lm.server --model "ikppramesh/irx-2" # Calling the OpenAI-compatible server with curl curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "ikppramesh/irx-2", "messages": [ {"role": "user", "content": "Hello"} ] }' - Hermes Agent
How to use ikppramesh/irx-2 with Hermes Agent:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "ikppramesh/irx-2"
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default ikppramesh/irx-2
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use ikppramesh/irx-2 with OpenClaw:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "ikppramesh/irx-2"
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "ikppramesh/irx-2" \ --custom-provider-id mlx-lm \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
IRx-2 — ⭐ Best Overall
The recommended model of the IRx family: a private, offline AI chat assistant with the best balance of answer quality and speed. Everything runs on your device — no internet connection, no account, and nothing you type ever leaves your phone, tablet or laptop.
The IRx family
All three run fully offline — nothing you type ever leaves your device.
| Model | Download | Best for | Phones & tablets (GGUF) | Mac (MLX) |
|---|---|---|---|---|
| IRx-1 | ~1.2GB | The fastest. Quick everyday answers on almost any phone | irx-1-GGUF | irx-1 |
| ⭐ IRx-2 — Best Overall | ~2.7GB | The best balance of answer quality and speed | irx-2-GGUF | irx-2 |
| IRx-2 Pro | ~5.3GB | The most detailed, long-form answers and long conversations | irx-2-pro-GGUF | irx-2-pro |
Not sure which to pick? Start with ⭐ IRx-2. Choose IRx-1 for speed or an older / smaller phone, and IRx-2 Pro for the richest answers on a device with 12GB+ RAM.
What IRx-2 does best — ⭐ Best Overall
IRx-2 is the recommended model for most people. It is about twice the size of IRx-1, which shows in clearer reasoning, better-structured answers, more reliable writing and stronger coding help — while still running comfortably on phones and tablets with 8GB+ RAM.
| Who | Example things to ask |
|---|---|
| Software developers | "Explain this error and how to fix it: …" · "Write a function that validates an email address, with tests." · "Why is this SQL query slow?" |
| Farmers | "Plan a monthly budget for a 2-acre vegetable farm." · "Compare drip and flood irrigation: pros and cons." · "Write a loan application letter to my bank." |
| Students | "Make a one-week study plan for my exams." · "Explain Newton's laws with everyday examples." |
| Teachers | "Create a 40-minute lesson plan on the water cycle, with a short quiz." |
| Shop & small business owners | "Outline a simple business plan for a tea stall." · "Write a polite reply to a customer complaint." |
| Writers & creators | "Outline a 5-minute YouTube script about saving money." · "Make this paragraph sound more professional." |
| Job seekers | "Improve these resume bullet points." · "Give me 10 likely interview questions for a sales role." |
| Everyday life | "Plan a family weekend on a budget." · "Help me write a birthday message for my father." |
Measured on the same Mac CPU: IRx-2 generates about 26 tokens/s vs IRx-1's 54 — roughly half the speed, for noticeably better answers.
Usage (Mac, MLX)
pip install mlx-lm
from mlx_lm import load, generate
from mlx_lm.sample_utils import make_sampler
model, tokenizer = load("ikppramesh/irx-2")
messages = [{"role": "user", "content": "How do I convert Celsius to Fahrenheit?"}]
prompt = tokenizer.apply_chat_template(messages, add_generation_prompt=True, tokenize=False)
print(generate(model, tokenizer, prompt=prompt, max_tokens=1024,
sampler=make_sampler(temp=0.7, top_p=0.95)))
No system prompt is required: the built-in chat template supplies the IRx-2 one when
none is given, and always keeps thinking mode off. Sample at a non-zero temperature
(e.g. temp=0.7). On a phone or tablet, use the
GGUF build.
Limitations
- Not a frontier-scale model — it won't match large hosted AI services on very hard multi-step reasoning, deep coding problems or breadth of world knowledge.
- Facts can be wrong. Like any small offline model it can state things confidently that aren't true — double-check anything important, such as prices, medicines and doses, laws, and local farming advice (seed varieties, chemical quantities, weather).
- Not a replacement for professionals — medical, legal and financial answers are general information only.
- Not current-events aware — its knowledge is fixed at training time.
- Don't enable native tool/function-calling in chat apps — never trained; plain chat is reliable.
Changelog
- 2026-09-30 — First release. Built with the IRx pipeline's fixes: fine-tune merged into the full-precision base and quantized once with an importance matrix (no repeating/looping replies), thinking mode always off, identity built in. Passed the automated pre-publish check: repeated 8-turn conversations with the repeat penalty off and the app's thinking request on — no loops, no copied answers, no hidden reasoning, correct identity.
License
Apache 2.0. IRx-2 is a derivative fine-tuned model — full Apache 2.0 terms apply as with any work under this license.
- Downloads last month
- 125
4-bit