File size: 2,282 Bytes
427445b
40a1398
427445b
 
 
02216b6
04a71c3
40a1398
427445b
40a1398
0728ee6
40a1398
0728ee6
40a1398
04a71c3
3db3124
40a1398
0728ee6
40a1398
ef4057b
40a1398
 
 
 
 
0728ee6
40a1398
0728ee6
40a1398
 
 
0728ee6
40a1398
0728ee6
 
 
 
 
40a1398
 
 
0728ee6
40a1398
0728ee6
40a1398
0728ee6
40a1398
1a53b9f
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
---
title: TensorFold
sdk: static
pinned: false
---

<div align="center">
  <img src="./tensorfold-hero.png" alt="TensorFold" width="100%">

# TensorFold

### Fast local LLM inference and tested model releases for Apple Silicon and NVIDIA CUDA.

[tensorfold.dev](https://tensorfold.dev) · [Runtime and setup](https://github.com/ashhart/TensorFold) · [Model collections](https://huggingface.co/TensorFold/collections)
</div>

TensorFold is an open-source inference runtime and a practical model-release project. The runtime serves local models behind an OpenAI-compatible API, while this Hugging Face organization publishes checkpoints and supporting assets tested on real hardware.

## What you will find here

- MLX quantized checkpoints for Apple Silicon
- MTP and DFlash assets where the upstream model provides a compatible drafter
- NVIDIA and DGX Spark recipes when a release has been tested there
- Measured speed, memory use, runtime versions, and known limits
- Clear credit, licences, and links to the original model authors

## Run with TensorFold

```bash
curl -fsSL https://tensorfold.dev/install.sh | sh
```

TensorFold supports macOS and Linux. See the [setup guide](https://github.com/ashhart/TensorFold) for current model families, runtime flags, and benchmark conditions.

## Choose a model for your Mac

| Unified memory | Collection | Selection basis |
| --- | --- | --- |
| 64 GB | [Browse models](https://huggingface.co/collections/TensorFold/mlx-models-for-64gb-macs) | Published peak below 48 GB |
| 128 GB | [Browse models](https://huggingface.co/collections/TensorFold/mlx-models-for-128gb-macs) | Published peak below 96 GB |
| 256 GB | [Browse models](https://huggingface.co/collections/TensorFold/mlx-models-for-256gb-macs) | Published peak below 192 GB |

These collections are starting points with nominal context headroom, not guarantees at maximum context. Each model card records the tested runtime, prompt, output length, memory evidence, and any compatibility caveats.

TensorFold does not train the base models. Model design, training, evaluations, and upstream documentation remain the work of the original authors and contributors.

[Follow TensorFold for new model releases, runtime updates, and fixes.](https://huggingface.co/TensorFold)