OPEN-SOURCE LOCAL INFERENCE
TensorFold
Fast local LLM inference and tested model releases for Apple Silicon and NVIDIA CUDA.
TensorFold is an open-source inference runtime with an OpenAI-compatible API. This Hugging Face organization publishes compatible checkpoints and supporting assets tested on real hardware.
What you will find here
- Apple Silicon: MLX quantized checkpoints with measured speed and memory use.
- Drafted decoding: MTP and DFlash assets when the upstream model provides a compatible drafter.
- NVIDIA CUDA: tested recipes for supported GPUs and DGX Spark systems.
- Clear release notes: runtime versions, setup steps, known limits, licences, and upstream credit.
Run TensorFold
curl -fsSL https://tensorfold.dev/install.sh | sh
TensorFold does not train the base models. Model design, training, evaluations, and upstream documentation remain the work of the original authors and contributors.