--- title: Spark-X2.5-4B Code Assistant emoji: 💻 colorFrom: indigo colorTo: blue sdk: gradio sdk_version: 6.15.1 app_file: app.py python_version: "3.12" startup_duration_timeout: 1h short_description: Streaming coding assistant powered by Spark-X2.5-4B. --- # 💻 Spark-X2.5-4B Code Assistant A streaming code-assistant demo for [XHToken/Spark-X2.5-4B](https://huggingface.co/XHToken/Spark-X2.5-4B), a compact general-purpose language model with strong coding, reasoning, agentic, and multilingual ability. The app runs the model directly in a Hugging Face **ZeroGPU** Space. It is tuned for programming tasks — write, explain, debug, refactor, test, and translate code — with task presets, a configurable system prompt, sampling controls, and an optional collapsible `` reasoning trace. ## Features - **Streaming responses** via `TextIteratorStreamer`, so code appears as it is generated. - **Reasoning trace** toggle — inspect Spark's `enable_thinking` chain-of-thought. - **Task presets** — general coding, write, explain, debug & fix, refactor, tests, translate. - **Markdown code rendering** with copy button, plus a copyable reasoning panel. - Exposed as an **MCP server** (`mcp_server=True`), so each handler is callable as a tool. ## Notes - The model is loaded with the official Transformers configuration, chat template, and custom `spark2_5` architecture (`trust_remote_code=True`) from its repository. - Inference uses `bfloat16` on the temporary ZeroGPU allocation, with the eager attention implementation required by the model's custom attention path. - Context is capped at 32,768 tokens here to keep interactive requests practical; the model itself supports a much larger native context window. - Defaults follow the model card: `temperature=1.0`, `top_p=0.95`, thinking on. ## License This demo code is Apache-2.0. The underlying model is released under its [Apache-2.0 license](https://huggingface.co/XHToken/Spark-X2.5-4B).