File size: 1,630 Bytes
e8cdc15
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
---
license: apache-2.0
base_model: Snowflake/Arctic-Text2SQL-R1-7B
tags:
- llamafile
- text-to-sql
- gguf
- qwen2
- reasoning
---

# Snowflake Arctic-Text2SQL-R1-7B - Llamafile

This repository provides a standalone, cross-platform executable [llamafile](https://github.com/mozilla-ai/llamafile) for [Snowflake/Arctic-Text2SQL-R1-7B](https://huggingface.co/Snowflake/Arctic-Text2SQL-R1-7B).

## What is a Llamafile?

A llamafile bundles the `llamafile` inference runtime together with model weights (`Q4_K_M` quantization) and default parameters into a single executable file. It runs locally on Linux, macOS, and Windows without needing Python, PyTorch, or package installations.

## Quick Start

### 1. Download & Make Executable

```bash
curl -L -o arctic-text2sql-r1-7b.llamafile https://huggingface.co/MigueldsBatista/Arctic-Text2SQL-R1-7B-llamafile/resolve/main/arctic-text2sql-r1-7b.llamafile
chmod +x arctic-text2sql-r1-7b.llamafile
```

### 2. Run Interactive / CLI Mode

```bash
./arctic-text2sql-r1-7b.llamafile --cli -p "<|im_start|>user
Write a SQL query to list all department names.
<|im_end|>
<|im_start|>assistant
"
```

### 3. Run Server Mode (OpenAI-compatible API)

```bash
./arctic-text2sql-r1-7b.llamafile --server --port 8080
```

Access the API at `http://localhost:8080/v1/chat/completions`.

## Hardware Acceleration Notes

- **GPU Acceleration**: Built-in CUDA acceleration works out-of-the-box on supported NVIDIA GPUs.
- **CPU Fallback / Blackwell GPUs**: For architectures not yet bundled in upstream llamafile (e.g. RTX 50-series sm_120) or systems without a dedicated GPU, append `--gpu disable`.