IRx-2 — ⭐ Best Overall

The recommended model of the IRx family: a private, offline AI chat assistant with the best balance of answer quality and speed. Everything runs on your device — no internet connection, no account, and nothing you type ever leaves your phone, tablet or laptop.

The IRx family

All three run fully offline — nothing you type ever leaves your device.

Model Download Best for Phones & tablets (GGUF) Mac (MLX)
IRx-1 ~1.2GB The fastest. Quick everyday answers on almost any phone irx-1-GGUF irx-1
⭐ IRx-2 — Best Overall ~2.7GB The best balance of answer quality and speed irx-2-GGUF irx-2
IRx-2 Pro ~5.3GB The most detailed, long-form answers and long conversations irx-2-pro-GGUF irx-2-pro

Not sure which to pick? Start with ⭐ IRx-2. Choose IRx-1 for speed or an older / smaller phone, and IRx-2 Pro for the richest answers on a device with 12GB+ RAM.

What IRx-2 does best — ⭐ Best Overall

IRx-2 is the recommended model for most people. It is about twice the size of IRx-1, which shows in clearer reasoning, better-structured answers, more reliable writing and stronger coding help — while still running comfortably on phones and tablets with 8GB+ RAM.

Who Example things to ask
Software developers "Explain this error and how to fix it: …" · "Write a function that validates an email address, with tests." · "Why is this SQL query slow?"
Farmers "Plan a monthly budget for a 2-acre vegetable farm." · "Compare drip and flood irrigation: pros and cons." · "Write a loan application letter to my bank."
Students "Make a one-week study plan for my exams." · "Explain Newton's laws with everyday examples."
Teachers "Create a 40-minute lesson plan on the water cycle, with a short quiz."
Shop & small business owners "Outline a simple business plan for a tea stall." · "Write a polite reply to a customer complaint."
Writers & creators "Outline a 5-minute YouTube script about saving money." · "Make this paragraph sound more professional."
Job seekers "Improve these resume bullet points." · "Give me 10 likely interview questions for a sales role."
Everyday life "Plan a family weekend on a budget." · "Help me write a birthday message for my father."

Measured on the same Mac CPU: IRx-2 generates about 26 tokens/s vs IRx-1's 54 — roughly half the speed, for noticeably better answers.

Usage (Mac, MLX)

pip install mlx-lm
from mlx_lm import load, generate
from mlx_lm.sample_utils import make_sampler

model, tokenizer = load("ikppramesh/irx-2")
messages = [{"role": "user", "content": "How do I convert Celsius to Fahrenheit?"}]
prompt = tokenizer.apply_chat_template(messages, add_generation_prompt=True, tokenize=False)
print(generate(model, tokenizer, prompt=prompt, max_tokens=1024,
               sampler=make_sampler(temp=0.7, top_p=0.95)))

No system prompt is required: the built-in chat template supplies the IRx-2 one when none is given, and always keeps thinking mode off. Sample at a non-zero temperature (e.g. temp=0.7). On a phone or tablet, use the GGUF build.

Limitations

  • Not a frontier-scale model — it won't match large hosted AI services on very hard multi-step reasoning, deep coding problems or breadth of world knowledge.
  • Facts can be wrong. Like any small offline model it can state things confidently that aren't true — double-check anything important, such as prices, medicines and doses, laws, and local farming advice (seed varieties, chemical quantities, weather).
  • Not a replacement for professionals — medical, legal and financial answers are general information only.
  • Not current-events aware — its knowledge is fixed at training time.
  • Don't enable native tool/function-calling in chat apps — never trained; plain chat is reliable.

Changelog

  • 2026-09-30 — First release. Built with the IRx pipeline's fixes: fine-tune merged into the full-precision base and quantized once with an importance matrix (no repeating/looping replies), thinking mode always off, identity built in. Passed the automated pre-publish check: repeated 8-turn conversations with the repeat penalty off and the app's thinking request on — no loops, no copied answers, no hidden reasoning, correct identity.

License

Apache 2.0. IRx-2 is a derivative fine-tuned model — full Apache 2.0 terms apply as with any work under this license.

Downloads last month
125
Safetensors
Model size
4B params
Tensor type
U32
·
BF16
·
F32
·
MLX
Hardware compatibility
Log In to add your hardware

4-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support