README / index.html
ashxhart's picture
Replace stale Vontra organization card
537c952 verified
Raw History Blame Contribute Delete
2.14 kB
<!doctype html>
<html lang="en">
<head>
<meta charset="utf-8" />
<meta name="viewport" content="width=device-width, initial-scale=1" />
<meta
name="description"
content="TensorFold provides fast local LLM inference and tested model releases for Apple Silicon and NVIDIA CUDA."
/>
<title>TensorFold | Local LLM inference</title>
</head>
<body>
<main>
<header>
<p><strong>OPEN-SOURCE LOCAL INFERENCE</strong></p>
<img
src="https://huggingface.co/spaces/TensorFold/README/resolve/main/tensorfold-logo.png"
alt="TensorFold"
width="120"
/>
<h1>TensorFold</h1>
<p><strong>Fast local LLM inference and tested model releases for Apple Silicon and NVIDIA CUDA.</strong></p>
<p>
<a href="https://tensorfold.dev">tensorfold.dev</a>
·
<a href="https://github.com/ashhart/TensorFold">GitHub</a>
</p>
</header>
<br />
<p>
TensorFold is an open-source inference runtime with an OpenAI-compatible API. This
Hugging Face organization publishes compatible checkpoints and supporting assets tested
on real hardware.
</p>
<br />
<h2>What you will find here</h2>
<ul>
<li><strong>Apple Silicon:</strong> MLX quantized checkpoints with measured speed and memory use.</li>
<li><strong>Drafted decoding:</strong> MTP and DFlash assets when the upstream model provides a compatible drafter.</li>
<li><strong>NVIDIA CUDA:</strong> tested recipes for supported GPUs and DGX Spark systems.</li>
<li><strong>Clear release notes:</strong> runtime versions, setup steps, known limits, licences, and upstream credit.</li>
</ul>
<br />
<h2>Run TensorFold</h2>
<p><code>curl -fsSL https://tensorfold.dev/install.sh | sh</code></p>
<br />
<p>
TensorFold does not train the base models. Model design, training, evaluations, and
upstream documentation remain the work of the original authors and contributors.
</p>
</main>
</body>
</html>