Spaces:
Running on Zero
Running on Zero
|
Download README.md from InfiniteDev/spark-x2.5-4b-code: direct link, hf CLI and curl.
- Browser
- Download file 1.98 kB
-
https://huggingface.co/spaces/InfiniteDev/spark-x2.5-4b-code/resolve/main/README.md
- Command line
-
hf download hf://spaces/InfiniteDev/spark-x2.5-4b-code/README.md
-
curl -L -o README.md https://huggingface.co/spaces/InfiniteDev/spark-x2.5-4b-code/resolve/main/README.md
1.98 kB
| title: Spark-X2.5-4B Code Assistant | |
| emoji: π» | |
| colorFrom: indigo | |
| colorTo: blue | |
| sdk: gradio | |
| sdk_version: 6.15.1 | |
| app_file: app.py | |
| python_version: "3.12" | |
| startup_duration_timeout: 1h | |
| short_description: Streaming coding assistant powered by Spark-X2.5-4B. | |
| # π» Spark-X2.5-4B Code Assistant | |
| A streaming code-assistant demo for | |
| [XHToken/Spark-X2.5-4B](https://huggingface.co/XHToken/Spark-X2.5-4B), a compact | |
| general-purpose language model with strong coding, reasoning, agentic, and | |
| multilingual ability. | |
| The app runs the model directly in a Hugging Face **ZeroGPU** Space. It is tuned | |
| for programming tasks β write, explain, debug, refactor, test, and translate code | |
| β with task presets, a configurable system prompt, sampling controls, and an | |
| optional collapsible `<think>` reasoning trace. | |
| ## Features | |
| - **Streaming responses** via `TextIteratorStreamer`, so code appears as it is generated. | |
| - **Reasoning trace** toggle β inspect Spark's `enable_thinking` chain-of-thought. | |
| - **Task presets** β general coding, write, explain, debug & fix, refactor, tests, translate. | |
| - **Markdown code rendering** with copy button, plus a copyable reasoning panel. | |
| - Exposed as an **MCP server** (`mcp_server=True`), so each handler is callable as a tool. | |
| ## Notes | |
| - The model is loaded with the official Transformers configuration, chat template, | |
| and custom `spark2_5` architecture (`trust_remote_code=True`) from its repository. | |
| - Inference uses `bfloat16` on the temporary ZeroGPU allocation, with the eager | |
| attention implementation required by the model's custom attention path. | |
| - Context is capped at 32,768 tokens here to keep interactive requests practical; | |
| the model itself supports a much larger native context window. | |
| - Defaults follow the model card: `temperature=1.0`, `top_p=0.95`, thinking on. | |
| ## License | |
| This demo code is Apache-2.0. The underlying model is released under its | |
| [Apache-2.0 license](https://huggingface.co/XHToken/Spark-X2.5-4B). | |