--- license: apache-2.0 base_model: Snowflake/Arctic-Text2SQL-R1-7B tags: - llamafile - text-to-sql - gguf - qwen2 - reasoning --- # Snowflake Arctic-Text2SQL-R1-7B - Llamafile This repository provides a standalone, cross-platform executable [llamafile](https://github.com/mozilla-ai/llamafile) for [Snowflake/Arctic-Text2SQL-R1-7B](https://huggingface.co/Snowflake/Arctic-Text2SQL-R1-7B). ## What is a Llamafile? A llamafile bundles the `llamafile` inference runtime together with model weights (`Q4_K_M` quantization) and default parameters into a single executable file. It runs locally on Linux, macOS, and Windows without needing Python, PyTorch, or package installations. ## Quick Start ### 1. Download & Make Executable ```bash curl -L -o arctic-text2sql-r1-7b.llamafile https://huggingface.co/MigueldsBatista/Arctic-Text2SQL-R1-7B-llamafile/resolve/main/arctic-text2sql-r1-7b.llamafile chmod +x arctic-text2sql-r1-7b.llamafile ``` ### 2. Run Interactive / CLI Mode ```bash ./arctic-text2sql-r1-7b.llamafile --cli -p "<|im_start|>user Write a SQL query to list all department names. <|im_end|> <|im_start|>assistant " ``` ### 3. Run Server Mode (OpenAI-compatible API) ```bash ./arctic-text2sql-r1-7b.llamafile --server --port 8080 ``` Access the API at `http://localhost:8080/v1/chat/completions`. ## Hardware Acceleration Notes - **GPU Acceleration**: Built-in CUDA acceleration works out-of-the-box on supported NVIDIA GPUs. - **CPU Fallback / Blackwell GPUs**: For architectures not yet bundled in upstream llamafile (e.g. RTX 50-series sm_120) or systems without a dedicated GPU, append `--gpu disable`.