offline-ai-models-nvidia

NVIDIA-only model catalog for the Offline AI Android app (models/manifest.json is merged into the app's picker automatically).

Every entry keeps url pointing at the upstream repo (NVIDIA official, bartowski, mradermacher, csukuangfj) β€” bytes download straight from the source, no re-uploads, and sha256 is the real LFS hash so the app's integrity check passes end to end.

Categories

type entries notes
llm Nemotron 3 Nano 4B β€” official Q4_K_M, IQ2_M, Q5_K_M nemotron_h hybrid (Mamba2 + 4 attn blocks); thinking-capable; ~2.9–3.6 GB RAM @8k ctx
llm Llama-3.1-Nemotron-Nano-4B-v1.1 Q4_K_M dense Llama arch, post-trained for RAG + tool calling
code OpenCodeReasoning-Nemotron-7B Q4_K_M NVIDIA's code-reasoning model (Qwen2 base); needs ~5.2 GB
stt Parakeet TDT 0.6B v2 int8 (sherpa-onnx bundle) NVIDIA NeMo ASR; enc+dec+joiner+tokens via parts
ocr β€” none β€” no small NVIDIA VL GGUF ships an mmproj; 8B/12B VL are too large for phones today
imgen β€” none β€” NVIDIA ships no sub-2 GB diffusion model
hq β€” pending β€” Zrald HQ-quantized Nemotron lands here after calibration

Adding an entry

Point url at any public HF repo, keep size/sha256 as the upstream LFS values, set min_ram_gb to file-size + KV estimate + ~400 MB. Multi-file bundles (STT/TTS) use parts with per-file file/size/ sha256/url.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support