File size: 2,491 Bytes
fc8ec75 ce30b2e 8104981 e623623 a5b93d6 e623623 2b41305 | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 | ---
title: README
emoji: π
colorFrom: red
colorTo: yellow
sdk: static
pinned: false
---
We are the InternSVG team from the Shanghai AI Laboratory, dedicated to empowering the InternVL series models with unified capabilities for SVG vector graphic understanding, editing, and generation.
Current Work:
## InternSVG: Towards Unified SVG Tasks with Multimodal Large Language Models
The InternSVG Family β a comprehensive suite that unifies data, benchmarks, and models for SVG understanding, editing, and generation. It consists of:
π§© SAgoge β the largest and most diverse multimodal SVG dataset, covering icons, illustrations, chemistry diagrams, and dynamic animations;
π SArena β a companion benchmark offering unified task definitions and standardized evaluation protocols across SVG domains;
π€ InternSVG Models β multimodal large language models trained for SVG understanding, editing, and generation.
Project Links
π Project Page: https://hmwang2002.github.io/release/internsvg/
π ArXiv Paper: https://arxiv.org/abs/2510.11341
π» GitHub Repository: https://github.com/hmwang2002/InternSVG
π SArena Benchmark: https://huggingface.co/datasets/InternSVG/SArena
π§© SAgoge Dataset: https://huggingface.co/datasets/InternSVG/SAgoge
π€ InternSVG-8B Model: https://huggingface.co/InternSVG/InternSVG-8B
## Reliable Reasoning in SVG-LLMs via Multi-Task Multi-Reward Reinforcement Learning
In this work, we present CTRL-S (Chain-of-Thought Reinforcement Learning for SVG), a unified framework that introduces a chain-of-thought mechanism to explicitly expose the modelβs reasoning process during SVG generation. To support this structured reasoning, we construct SVG-Sophia, a high-quality dataset of 145K samples across SVG code refinement, Text-to-SVG, and Image-to-SVG tasks. Furthermore, we design a robust multi-reward reinforcement learning scheme powered by the GRPO algorithm. By jointly optimizing across DINO, image-text similarity, format, and code efficiency rewards in a multi-task setting, our approach systematically boosts structural coherence and generation capabilities. Extensive experiments show that CTRL-S outperforms existing methods, achieving higher task success rates, superior code quality, and exceptional visual fidelity.
π ArXiv Paper: https://arxiv.org/abs/2603.16189
π» GitHub Repository: https://github.com/hmwang2002/CTRL-S
π§© SVG-Sophia Dataset: https://huggingface.co/datasets/InternSVG/SVG-Sophia
|