README / README.md
ashxhart's picture
Refresh TensorFold organization card
1a53b9f verified
|
Raw History Blame Contribute Delete
2.28 kB
metadata
title: TensorFold
sdk: static
pinned: false
TensorFold

TensorFold

Fast local LLM inference and tested model releases for Apple Silicon and NVIDIA CUDA.

tensorfold.dev · Runtime and setup · Model collections

TensorFold is an open-source inference runtime and a practical model-release project. The runtime serves local models behind an OpenAI-compatible API, while this Hugging Face organization publishes checkpoints and supporting assets tested on real hardware.

What you will find here

  • MLX quantized checkpoints for Apple Silicon
  • MTP and DFlash assets where the upstream model provides a compatible drafter
  • NVIDIA and DGX Spark recipes when a release has been tested there
  • Measured speed, memory use, runtime versions, and known limits
  • Clear credit, licences, and links to the original model authors

Run with TensorFold

curl -fsSL https://tensorfold.dev/install.sh | sh

TensorFold supports macOS and Linux. See the setup guide for current model families, runtime flags, and benchmark conditions.

Choose a model for your Mac

Unified memory Collection Selection basis
64 GB Browse models Published peak below 48 GB
128 GB Browse models Published peak below 96 GB
256 GB Browse models Published peak below 192 GB

These collections are starting points with nominal context headroom, not guarantees at maximum context. Each model card records the tested runtime, prompt, output length, memory evidence, and any compatibility caveats.

TensorFold does not train the base models. Model design, training, evaluations, and upstream documentation remain the work of the original authors and contributors.

Follow TensorFold for new model releases, runtime updates, and fixes.