Instructions to use ikppramesh/irx-2-pro with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- MLX
How to use ikppramesh/irx-2-pro with MLX:
# Make sure mlx-lm is installed # pip install --upgrade mlx-lm # Generate text with mlx-lm from mlx_lm import load, generate model, tokenizer = load("ikppramesh/irx-2-pro") prompt = "Write a story about Einstein" messages = [{"role": "user", "content": prompt}] prompt = tokenizer.apply_chat_template( messages, add_generation_prompt=True ) text = generate(model, tokenizer, prompt=prompt, verbose=True) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- LM Studio
- Pi
How to use ikppramesh/irx-2-pro with Pi:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "ikppramesh/irx-2-pro"
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "mlx-lm": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "ikppramesh/irx-2-pro" } ] } } }Run Pi
# Start Pi in your project directory: pi
- MLX LM
How to use ikppramesh/irx-2-pro with MLX LM:
Generate or start a chat session
# Install MLX LM uv tool install mlx-lm # Interactive chat REPL mlx_lm.chat --model "ikppramesh/irx-2-pro"
Run an OpenAI-compatible server
# Install MLX LM uv tool install mlx-lm # Start the server mlx_lm.server --model "ikppramesh/irx-2-pro" # Calling the OpenAI-compatible server with curl curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "ikppramesh/irx-2-pro", "messages": [ {"role": "user", "content": "Hello"} ] }' - Hermes Agent
How to use ikppramesh/irx-2-pro with Hermes Agent:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "ikppramesh/irx-2-pro"
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default ikppramesh/irx-2-pro
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use ikppramesh/irx-2-pro with OpenClaw:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "ikppramesh/irx-2-pro"
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "ikppramesh/irx-2-pro" \ --custom-provider-id mlx-lm \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
IRx-2 Pro
The largest edition of the IRx family: a private, offline AI chat assistant for long, detailed answers and extended conversations. Everything runs on your device — no internet connection, no account, and nothing you type ever leaves it.
The IRx family
All three run fully offline — nothing you type ever leaves your device.
| Model | Download | Best for | Phones & tablets (GGUF) | Mac (MLX) |
|---|---|---|---|---|
| IRx-1 | ~1.2GB | The fastest. Quick everyday answers on almost any phone | irx-1-GGUF | irx-1 |
| ⭐ IRx-2 — Best Overall | ~2.7GB | The best balance of answer quality and speed | irx-2-GGUF | irx-2 |
| IRx-2 Pro | ~5.3GB | The most detailed, long-form answers and long conversations | irx-2-pro-GGUF | irx-2-pro |
Not sure which to pick? Start with ⭐ IRx-2. Choose IRx-1 for speed or an older / smaller phone, and IRx-2 Pro for the richest answers on a device with 12GB+ RAM.
What IRx-2 Pro does best
IRx-2 Pro is the largest edition. It was trained with the longest context window of the family (so it learned from nearly all of the long conversations in the training data), which makes it the best choice for long, detailed answers and for working through a topic over many messages. It runs at about the same speed as IRx-2, but needs more memory: use it on devices with 12GB+ RAM, or on a Mac.
| Who | Example things to ask |
|---|---|
| Software developers | "Help me design a small inventory app step by step — data model, screens and API." · "Walk me through refactoring this code over a few messages." |
| Farmers | "Help me write a full-season plan for my farm — sowing, watering, spraying and harvest — then let's refine it." · "Write a detailed proposal to sell my produce to a local supermarket." |
| Students & researchers | "Turn these notes into a structured report with headings and a summary." · "Explain this topic in depth, then quiz me on it." |
| Teachers | "Create a full week's unit on fractions: daily lessons, activities and a final test." |
| Business owners | "Draft a detailed marketing plan for my new bakery." · "Write a complete employee handbook outline." |
| Writers & creators | "Write a 1,500-word blog post on healthy habits." · "Help me develop a short story over several messages." |
| Everyday life | "Plan a 7-day family trip with a day-by-day schedule and budget." |
Measured on the same Mac CPU: IRx-2 Pro generates about 24 tokens/s, close to IRx-2's 26 — most of its extra size is lookup tables that cost little compute.
Usage (Mac, MLX)
pip install mlx-lm
from mlx_lm import load, generate
from mlx_lm.sample_utils import make_sampler
model, tokenizer = load("ikppramesh/irx-2-pro")
messages = [{"role": "user", "content": "How do I convert Celsius to Fahrenheit?"}]
prompt = tokenizer.apply_chat_template(messages, add_generation_prompt=True, tokenize=False)
print(generate(model, tokenizer, prompt=prompt, max_tokens=1024,
sampler=make_sampler(temp=0.7, top_p=0.95)))
No system prompt is required: the built-in chat template supplies the IRx-2 Pro one when
none is given, and always keeps thinking mode off. Sample at a non-zero temperature
(e.g. temp=0.7). On a phone or tablet, use the
GGUF build.
Limitations
- Not a frontier-scale model — it won't match large hosted AI services on very hard multi-step reasoning, deep coding problems or breadth of world knowledge.
- Facts can be wrong. Like any small offline model it can state things confidently that aren't true — double-check anything important, such as prices, medicines and doses, laws, and local farming advice (seed varieties, chemical quantities, weather).
- Not a replacement for professionals — medical, legal and financial answers are general information only.
- Not current-events aware — its knowledge is fixed at training time.
- Don't enable native tool/function-calling in chat apps — never trained; plain chat is reliable.
Changelog
- 2026-09-30 — First release (published earlier the same day under a different repository name; old links redirect here). Trained with a 2048-token window, merged into the full-precision base and quantized once with an importance matrix. Passed the automated pre-publish check: repeated 8-turn conversations with the repeat penalty off and the app's thinking request on — no loops, no copied answers, no hidden reasoning, correct identity.
License
Apache 2.0. IRx-2 Pro is a derivative fine-tuned model — full Apache 2.0 terms apply as with any work under this license.
- Downloads last month
- 27
4-bit