DeepWiki Coder 7B v2 — MLX 4-bit

What?

This is the MLX version of DeepWiki Coder 7B v2. It is a 4-bit quantized language model for Apple Silicon that generates technical documentation from source-code context.

Why?

MLX is optimized for Apple Silicon's unified memory architecture. 4-bit quantization substantially reduces memory use, making this 7B model practical on consumer Macs while trading away some quality compared with the merged full-precision model.

Quick start

pip install -U mlx-lm
mlx_lm.generate \
  --model GhostScientist/semanticwiki-coder-7b-v2-mlx-4bit \
  --prompt "Explain this codebase architecture."

Or from Python:

from mlx_lm import generate, load

model, tokenizer = load("GhostScientist/semanticwiki-coder-7b-v2-mlx-4bit")
print(generate(model, tokenizer, prompt="Explain this codebase architecture."))

Limitations

Quantization can reduce accuracy and citation reliability. Review generated documentation, especially when source context is incomplete or very large. The model is not a security auditor, and prompts should not contain secrets.

Provenance

Downloads last month
28
Safetensors
Model size
8B params
Tensor type
U32
·
BF16
·
MLX
Hardware compatibility
Log In to add your hardware

4-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for GhostScientist/semanticwiki-coder-7b-v2-mlx-4bit