MigueldsBatista's picture
Upload README.md with huggingface_hub
e8cdc15 verified
|
Raw History Blame Contribute Delete
1.63 kB
---
license: apache-2.0
base_model: Snowflake/Arctic-Text2SQL-R1-7B
tags:
- llamafile
- text-to-sql
- gguf
- qwen2
- reasoning
---
# Snowflake Arctic-Text2SQL-R1-7B - Llamafile
This repository provides a standalone, cross-platform executable [llamafile](https://github.com/mozilla-ai/llamafile) for [Snowflake/Arctic-Text2SQL-R1-7B](https://huggingface.co/Snowflake/Arctic-Text2SQL-R1-7B).
## What is a Llamafile?
A llamafile bundles the `llamafile` inference runtime together with model weights (`Q4_K_M` quantization) and default parameters into a single executable file. It runs locally on Linux, macOS, and Windows without needing Python, PyTorch, or package installations.
## Quick Start
### 1. Download & Make Executable
```bash
curl -L -o arctic-text2sql-r1-7b.llamafile https://huggingface.co/MigueldsBatista/Arctic-Text2SQL-R1-7B-llamafile/resolve/main/arctic-text2sql-r1-7b.llamafile
chmod +x arctic-text2sql-r1-7b.llamafile
```
### 2. Run Interactive / CLI Mode
```bash
./arctic-text2sql-r1-7b.llamafile --cli -p "<|im_start|>user
Write a SQL query to list all department names.
<|im_end|>
<|im_start|>assistant
"
```
### 3. Run Server Mode (OpenAI-compatible API)
```bash
./arctic-text2sql-r1-7b.llamafile --server --port 8080
```
Access the API at `http://localhost:8080/v1/chat/completions`.
## Hardware Acceleration Notes
- **GPU Acceleration**: Built-in CUDA acceleration works out-of-the-box on supported NVIDIA GPUs.
- **CPU Fallback / Blackwell GPUs**: For architectures not yet bundled in upstream llamafile (e.g. RTX 50-series sm_120) or systems without a dedicated GPU, append `--gpu disable`.