offline-ai-models-nvidia
NVIDIA-only model catalog for the Offline AI Android app
(models/manifest.json is merged into the app's picker automatically).
Every entry keeps url pointing at the upstream repo (NVIDIA official,
bartowski, mradermacher, csukuangfj) β bytes download straight from the
source, no re-uploads, and sha256 is the real LFS hash so the app's
integrity check passes end to end.
Categories
| type | entries | notes |
|---|---|---|
llm |
Nemotron 3 Nano 4B β official Q4_K_M, IQ2_M, Q5_K_M | nemotron_h hybrid (Mamba2 + 4 attn blocks); thinking-capable; ~2.9β3.6 GB RAM @8k ctx |
llm |
Llama-3.1-Nemotron-Nano-4B-v1.1 Q4_K_M | dense Llama arch, post-trained for RAG + tool calling |
code |
OpenCodeReasoning-Nemotron-7B Q4_K_M | NVIDIA's code-reasoning model (Qwen2 base); needs ~5.2 GB |
stt |
Parakeet TDT 0.6B v2 int8 (sherpa-onnx bundle) | NVIDIA NeMo ASR; enc+dec+joiner+tokens via parts |
ocr |
β none β | no small NVIDIA VL GGUF ships an mmproj; 8B/12B VL are too large for phones today |
imgen |
β none β | NVIDIA ships no sub-2 GB diffusion model |
hq |
β pending β | Zrald HQ-quantized Nemotron lands here after calibration |
Adding an entry
Point url at any public HF repo, keep size/sha256 as the upstream
LFS values, set min_ram_gb to file-size + KV estimate + ~400 MB.
Multi-file bundles (STT/TTS) use parts with per-file file/size/
sha256/url.
Inference Providers NEW
This model isn't deployed by any Inference Provider. π Ask for provider support