SenseNova-U1-8B-MoT-Infographic-V3

English | 简体中文

arXiv HuggingFace Model ModelScope-模型 SenseNova-U1 Demo License Discord

SenseNova-U1

📣 Updated News

  • [2026.07.16] Release SenseNova-U1-8B-MoT-Infographic-V3 📊, designed for integrated infographic generation and editing. It retains strong text-to-image (T2I) capabilities while significantly enhancing infographic editing, supporting localized text and content editing, global style editing, and global layout editing. See ✨ U1 Infographic Model Series for model details and benchmark results.
✨ Click to expand older news

🌟 Overview

🚀 SenseNova U1 is a new series of native multimodal models that unifies multimodal understanding, reasoning, and generation within a monolithic architecture. It marks a fundamental paradigm shift in multimodal AI: from modality integration to true unification. Rather than relying on adapters to translate between modalities, SenseNova U1 models think-and-act across language and vision natively.

✨ Click to expand architecture details

Unifying visual understanding and generation in an end-to-end architecture from pixel to word opens tremendous possibilities, enabling highly efficient and strong understanding, generation, and interleaved reasoning in a natively multimodal manner.

radar plot

🏗️ Key Pillars:

At the core of SenseNova U1 is NEO-unify, a novel architecture designed from the first principles for multimodal AI: It eliminates both Visual Encoder (VE) and Variational Auto-Encoder (VAE) where pixel-word information are inherently and deeply correlated. Several important features are as follows:

  • 🔗 Model language and visual information end-to-end as a unified compound.
  • 🖼️ Preserve semantic richness while maintaining pixel-level visual fidelity.
  • 🧠 Reason across modalities with high efficiency & minimal conflict via native MoTs.

Powered by this new core architecture, SenseNova-U1-8B-MoT-Infographic-V3 is the recommended release for an integrated infographic generation-and-editing workflow. It retains strong text-to-image (T2I) capability while adding image editing (IT2I) within the same unified model.

SenseNova-U1-8B-MoT-Infographic-V3 Image Editing Benchmark Performance

  • Benchmark Performance: V3 reaches an Overall Total of 50.23 on Qwen-Image-Bench, improving by 2.23 over V2's 48.00. It also achieves an Average of 5.89 on WeEdit and an overall score of 7.89 on GEdit-Bench, ranking first among the open-source models included in both editing evaluations.
  • Generation and Editing Quality: V3 supports both region-marked precise changes and natural-language-only editing. It handles localized text edits, local content insertion, removal, and replacement, as well as global style and layout editing, while preserving unedited regions whenever possible.
✨ Click to expand Benchmark Details

WeEdit (Image Editing)

Model Instruction Adherence ↑ Text Clarity ↑ Background Preservation ↑ Average ↑
Closed-source Models
Nano-Banana-Pro 8.58 9.10 8.85 8.84
Seedream 4.5 6.29 7.66 6.38 6.78
Gemini-2.5-Flash-Image 3.92 7.14 7.80 6.29
Qwen-Image-2.0 5.04 6.08 5.68 5.60
Open-source Models
SenseNova-U1-8B-MoT-Infographic-V3 5.67 5.94 6.06 5.89
FireRed-Image-Edit 4.15 6.33 7.14 5.87
HY-Image-3-Instruct 4.16 5.99 7.03 5.73
Qwen-Image-Edit-2509 3.49 5.84 6.80 5.38
LongCat-Image-Edit 2.78 4.36 7.03 4.72
Qwen-Image-Edit-2511 3.18 3.93 4.63 3.91
BAGEL 1.97 4.01 5.75 3.91

Note: Higher is better. Group-best scores are bolded. Average is the arithmetic mean of the three metrics. V3 ranks first by Average among the open-source models included in this evaluation.

GEdit-Bench (Image Editing)

Model EN_G_SC ↑ EN_G_PQ ↑ EN_G_O ↑
Closed-source Models
Qwen-Image-2.0 9.02 8.02 8.37
UniWorld-V2 8.39 8.02 7.83
Seedream 4.5 8.27 8.17 7.82
Nano-Banana-Pro 8.10 8.34 7.74
Seedream 4.0 8.14 8.12 7.70
Nano-Banana 7.40 8.45 7.29
Open-source Models
SenseNova-U1-8B-MoT-Infographic-V3 8.65 7.58 7.89
LongCat-Image-Edit 8.18 8.00 7.64
Qwen-Image-Edit-2511 8.00 7.86 7.56
Qwen-Image-Edit-2509 8.15 7.86 7.54
Step1X-Edit 7.66 7.35 6.97
BAGEL 7.36 6.83 6.52
OmniGen2 7.16 6.77 6.41
UniWorld-v1 4.93 7.43 4.85
AnyEdit 3.18 5.82 3.21

Note: Higher is better. Group-best scores are bolded. V3 delivers the best overall performance among the open-source models included in this evaluation.

Qwen-Image-Bench (Text-to-Image)

Model Quality ↑ Aesthetics ↑ Alignment ↑ Real-world Fidelity ↑ Creative Generation ↑ Overall Total ↑
GPT-Image-2 59.09 68.48 65.78 59.40 75.34 64.58
Nano-Banana-Pro 55.30 61.38 60.30 55.91 64.54 57.84
Qwen-Image-2.0 55.16 60.36 57.86 53.06 63.59 57.84
HunyuanImage-3.0 50.76 54.66 53.16 45.33 48.33 50.81
SenseNova-U1-8B-MoT-Infographic-V3 49.53 52.82 52.00 44.32 48.34 50.23
SenseNova-U1-8B-MoT-Infographic-V2 48.15 49.60 50.08 43.68 44.65 48.00
SenseNova-U1-8B-MoT-Infographic 47.12 48.15 48.90 43.35 45.40 47.11
HiDream-O1 44.20 45.35 43.74 40.28 36.24 42.84

Note: Higher is better. Column-best scores and the V3 model name are bolded. V3 reaches an Overall Total of 50.23, improving by 2.23 over V2's 48.00.

Historical Infographic Generation Benchmarks

Expand for complete BizGenEval, IGenBench, and OneIG results
Model BizGenEval Avg. (hard / easy) ↑ IGenBench Q-ACC ↑ IGenBench I-ACC ↑ OneIG(EN) ↑ OneIG(ZH) ↑
Commercial Models
Nano-Banana-Pro 76.7 / 93.7 90.6 48.8 58.1 56.8
Nano-Banana-2.0 68.5 / 92.5 85.6 34.4 54.0 54.9
GPT-Image-1.5 35.9 / 81.6 55.0 12.0 - -
Qwen-Image-2.0 45.5 / 65.8 50.0 3.0 54.1 50.9
Seedream-4.5 30.1 / 66.2 61.0 6.0 56.4 55.0
Open-source Models
SenseNova-U1-8B-MoT-Infographic-V3 49.5 / 68.5 66.3 13.2 54.6 52.4
SenseNova-U1-8B-MoT-Infographic-V2 50.3 / 67.9 71.4 18.3 55.4 53.5
SenseNova-U1-8B-MoT-Infographic 46.6 / 65.4 69.5 17.0 55.6 53.3
SenseNova-U1-8B-MoT 39.8 / 61.1 51.3 4.2 54.5 53.8
Z-Image 8.2 / 43.8 30.0 1.0 54.6 53.5
Qwen-Image-2512 6.3 / 41.0 32.2 1.0 53.0 51.5
Qwen-Image 2.8 / 23.8 36.0 0.0 53.9 54.8
Bagel 2.0 / 3.7 4.9 0.0 36.1 37.0

Note: V3 is the current release and is pinned to the first row of the open-source section. IGenBench scores are reported as percentages; the remaining commercial and open-source models are sorted separately by the average of BizGenEval hard/easy and IGenBench Q-ACC/I-ACC. OneIG is included as a general generation reference.

V2 and V3 are independently trained from the same base model with different task and data mixtures. V3 jointly trains infographic text-to-image generation and image editing with a lower proportion of IGenBench-related training data than V2, resulting in lower IGenBench scores than the generation-focused V2. At the same time, V3 remains broadly comparable to V2 on BizGenEval, improves the broader Qwen-Image-Bench score from 48.00 to 50.23, and adds infographic image-editing capability.

## SenseNova-U1-8B-MoT-Infographic-V3 Showcase

The following V3 examples are organized around integrated infographic generation-and-editing capabilities. They cover T2I generation and IT2I editing scenarios, including four major editing capability groups: local text editing, local content editing, global style editing, and global layout editing, and demonstrate precise dense-text repair and consistency preservation in non-edited regions.

Local Text Editing

Region-Marked Precise Editing

Instruction Before Edit After Edit
Instruction
将蓝色框中的屋顶变为红色,将红框中的标题改为比较简短的格式,将绿框中的屋子改为红色屋顶的咖啡屋,将紫色框中改为共享自行车站,将粉色框中的内容改为书柜,将右下角青色框中的内容改为,室内植物玻璃房
V3 edit showcase 26 input V3 edit showcase 26 output
Instruction
把红色框的“经典·从容·每一刻”换成经典时尚手表,蓝色框换成“100米防水”,橙色框换成电池寿命约10年
V3 edit showcase 03 input V3 edit showcase 03 output
Instruction
修复红框内模糊标题为“量子互联网的第一条链路”,去掉红框,其余不动。
V3 edit showcase 04 input V3 edit showcase 04 output
Instruction
请在图像中红色边界框标出的空白区域内插入指定内容,并使新增内容与真实环境的材质、透视、光照和空间关系自然融合。
目标区域:
“画面左下角主面板下方的空白深色区域,位于左侧四个数据卡片下方、底部图表左侧的红色边界框内部”
需要插入的文字或内容为:
“年度结余目标:120,000元”
请严格将新增内容限制在红框内部。完成插入后,需要彻底去除红色边界框,使最终图像中不再显示任何红线、控制点、箭头或编辑标记。
在插入内容之前,请先分析目标区域的实际环境,包括:
深色金融仪表盘背景的平面方向;
目标区域的矩形仪表盘面板边界;
玻璃拟态卡片、深色渐变背景和细线边框材质;
光源方向和亮度;
环境阴影与微弱高光;
字体应该像金融仪表盘中的数字标签和状态卡片,而不是普通贴纸文字。
新增文字必须按照目标表面的透视自然放置。由于目标区域是正面仪表盘平面,文字应保持水平排版、居中或左对齐,并与原有卡片网格对齐。
新增内容的颜色、亮度和清晰度必须符合环境光照。建议使用浅色文字配合绿色或金色强调数字,使其与原图“储蓄率 41%”“目标完成 76%”等数据卡片风格一致。不能使用与场景完全无关的纯白贴字,也不能让文字看起来悬浮在表面前方。
如果目标区域是仪表盘卡片:
文字应在红框内部合理居中;
保留原有深色面板边缘、微弱发光和阴影;
不得覆盖相邻卡片、图表或边框;
可以根据原有设计加入适当内边距和小型图标。
插入的文字必须完整、清晰、无乱码、无错别字。字体风格应与金融仪表盘和数据卡片用途相符,不得使用与场景冲突的字体。
除红框内部区域和红框本身之外,图像中的其他内容必须保持完全不变。不得移动标题“个人年度财务仪表盘”、左侧数据卡、现金流图、储蓄仪表、支出分类图、目标追踪列表或任何其他物体,不得改变整体光照、颜色和构图。
不得扩大目标区域,不得把新增内容放到红框外,也不得保留红框作为最终设计元素。
最重要的是,新增内容必须像原本就存在于该金融仪表盘中,而不是后来粘贴上去。最终画面应具有正确的材质融合、环境光照、卡片层级和自然对齐关系。
V3 edit showcase 05 input V3 edit showcase 05 output
Instruction
将原西班牙语标题替换为英文“STOP USING SHOWER SPONGES!”和副标题“YOUR SKIN DESERVES BETTER”,移除红框,框外所有文字、人物、颜色和布局完全不变。
V3 edit showcase 06 input V3 edit showcase 06 output
Instruction
翻译红框中的内容
V3 edit showcase 07 input V3 edit showcase 07 output
Instruction
把红色框的“我不见,黄河之水地下来”替换成“君不见,黄河之水天上来”,把蓝色框的“我不见,高堂明镜悲白发”替换成“君不见,高堂明镜悲白发”。
V3 edit showcase 08 input V3 edit showcase 08 output

Natural-Language Prompt Editing

Instruction Before Edit After Edit
Instruction
主题风格换成漫画风格,需要进行以下修改:
第一项修改:将原图中的文字:“STYLE”替换为:“时尚杂志”
第二项修改:将原图中的文字:“FORWARD”替换为:“美丽的补充力”
V3 edit showcase 01 input V3 edit showcase 01 output
Instruction

第一项修改:将原图中的文字:“不够高级“替换为:“DESIGN!OR DIE!”第二项修改:将原图中的文字:“没有惊喜感”,换成“缺乏设计感会导致”
V3 edit showcase 09 input V3 edit showcase 09 output
Instruction

第一项修改:将原图中的文字:“城市低碳生活地图”替换为:“城市绿色出行地图”第二项修改:将原图中的文字:“绿色出行 42%”替换为:“绿色出行 56%”第三项修改:将位于画面右上角卡片中的“节能建筑 36栋”,调整为“低碳建筑 48栋”。第四项修改:删除文字“社区回收 18站”中的“社区”,保留为“回收 18站”,并确保删除后语句自然、排版完整。第五项修改:修正原图中的数据:“公共绿地 128公顷”修改为“公共绿地 156公顷”。
V3 edit showcase 10 input V3 edit showcase 10 output
Instruction
把“主要航海壮举”换成“主要事件”,把“地理大发现时代”换成“!!地理大发现时期!!”,把“克里斯托弗·哥伦布”,换成“Cristóbal Colón”
V3 edit showcase 18 input V3 edit showcase 18 output
Instruction
把“人类进化史”换成“人类——我们的故事”,把“概述”换成“总结”,把“主要进化阶段”换成“演化路径”。
V3 edit showcase 19 input V3 edit showcase 19 output

Local Content Editing

Instruction Before Edit After Edit
Instruction
删除海面下的塑料袋、塑料瓶和吸管,将中央口号改为白色手写体“SAVE OUR OCEAN”,把右侧海龟移到文字下方并在底部新增回收标志,保持海洋分层结构清晰。
V3 edit showcase 02 input V3 edit showcase 02 output
Instruction
透明杯壁上加入“CURLY & PROUD”,文字需随杯面弧度自然变形,并呈现透过玻璃与茶水后的透明度和色偏。
V3 edit showcase 11 input V3 edit showcase 11 output
Instruction
复古电脑屏幕内加入绿色单色像素字“DESIGN MODE: ON”,匹配屏幕透视、颗粒噪点与玻璃反光;
V3 edit showcase 12 input V3 edit showcase 12 output
Instruction
将四代员工使用个人AI工具的比例依次改为:Gen Z 92%、Millennials 81%、Gen X 68%、Boomers 55%,并同步调整四根彩色柱子的高度。保留人物插画、英文标题、配色纹理、来源和品牌信息不变。
V3 edit showcase 13 input V3 edit showcase 13 output
Instruction
在顶部蓝色标题区右上角自然加入一枚金黄色奖杯徽章,奖杯内写“TOP GROWTH”,采用与原图一致的扁平矢量风格、粗线条和蓝黄配色,并添加少量庆祝星光装饰。调整徽章大小避免遮挡“2023 HIGHLIGHTS”,其他文字、数据、插画和布局保持不变。
V3 edit showcase 16 input V3 edit showcase 16 output

Global Style Editing

Instruction Before Edit After Edit
Instruction
主题风格换成乐高风格。将原图中的文字:“2022“替换为:“2025”第二项修改:将原图中的文字:“FINESTBUSINESSTRENDS”换成“HELLO! WORLD!”
V3 edit showcase 17 input V3 edit showcase 17 output
Instruction
把图片风格换成乐高风格和中国春节风格,所有文字换成像素字体。
V3 edit showcase 20 input V3 edit showcase 20 output
Instruction
将顶部主标题“城市应急避险指挥图”替换为“城市防汛应急联动图”。风格换成赛博朋克风格。
V3 edit showcase 21 input V3 edit showcase 21 output
Instruction
把图片风格换成中国传统风格。字体换成宋体
V3 edit showcase 22 input V3 edit showcase 22 output
Instruction
保持内容不变,迁移为亮色风格
V3 edit showcase 23 input V3 edit showcase 23 output
Instruction
风格换成复古羊皮纸地图风;保留全部文字、数据与版式。
V3 edit showcase 24 input V3 edit showcase 24 output

Global Layout Editing

Instruction Before Edit After Edit
Instruction
信息大致不变,对布局进行美化
V3 edit showcase 25 input V3 edit showcase 25 output

🛠️ Quick Start

🌐 Use with SenseNova-Studio

The fastest way to experience SenseNova-U1 is through SenseNova-Studio — a 🆓 free online playground where you can try the model directly in your browser, no installation or GPU required.

Note: To serve more users, U1-Fast has undergone step and CFG distillation, and is dedicated to infographic generation.

🦞 Use with SenseNova-Skills (OpenClaw)

The easiest way to integrate SenseNova-U1 into your own agent or application is through our companion repository SenseNova-Skills (OpenClaw) 🦞, which ships SenseNova-U1 as a ready-to-use skill with a unified tool-calling interface.

Refer to the SenseNova-Skills README for installation and usage details.

✨ Click to collapse and view interesting cases made through Skills and Studio

Skill Cases

🤗 Run with transformers (Default)

Setup: Follow the Installation Guide to clone the repo and install dependencies with uv.

🌟 Generate High-Quality Infographics

For generating complex infographics, we highly recommend using the following parameters: --cfg_scale 4.0, --timestep_shift 3.0, and --num_steps 50.

python examples/t2i/inference.py \
  --model_path sensenova/SenseNova-U1-8B-MoT-Infographic-V3 \
  --prompt "该信息图以复古漫画/动漫风格为主题设计,整体背景为深蓝色星空背景(带有颗粒质感),由多个带有技术网格线和星座线图案装饰的模块围绕中心主题呈海报式上下左右分布布局。中心是一个由两位角色与驾驶舱构成的主体视觉区域,内部包含主标题“SENSETIME”(带星星装饰的黄色大号粗体字),副标题“2026 走向未来!”(蓝色文字),下方有左上角信息框和文字“任务编号:K-2026-ALPHA | 发布日期:2026.06.27 | 保密等级:公开”。中心上方有散落的黄色星星与星座线图案装饰。 信息图共分为六个主要部分,编号1至6,分布在中心周围: **1. 顶部区域(标题与信息框模块)** - **主标题**(黄色大号粗体字带星星装饰):“SENSETIME”。 - **副标题**(蓝色文字):“2026 走向未来!”。 - **左上角信息框**(信息标签):包含“任务编号:K-2026-ALPHA”、“发布日期:2026.06.27”、“保密等级:公开”。 **2. 中心主视觉(角色与装备展示模块)** 以三位角色及驾驶舱按顺序排列,标签连接: - **左侧男性角色**(戴护目镜、红围巾、弹药带,手持武器,帅气姿态动作):标签文字“左轮手枪”(指向枪套),描述文字“标准配备|口径:.45|弹容量:6发|重量:1.1kg”。 - **中心女性角色**(戴护目镜、红围巾、条纹衫,弹奏白色空心电吉他,吉他上有“KATAO”字样,优美姿态):标签文字“格雷奇”(指向角色),描述文字“装备型号:G-77电吉他|频率范围:80Hz-12kHz|输出功率:50W”。 - **背景右侧角色**(驾驶舱内两名角色,女性调整护目镜,男性戴护目镜看前方):驾驶舱标签“驾驶舱控制面板|导航系统:GPS-2025|通讯频道:CH-77”。 **3. 左侧边栏(任务简报模块)** 包含一段说明文字:“任务简报” 下方配有五项核心指标列表: - 目标:探索未知领域 - 时间:2025年启动 - 地点:黑空 - 队员:精英组 - 装备:未来科技武装 **4. 右侧边栏(成员档案模块)** 包含一段说明文字:“探险队成员档案” 下方配有五名成员及技能列表: - 队长:经验丰富的领航员 - 副队长:技术专家 - 成员A:武器专家 - 成员B:通讯专员 - 特殊技能:时空导航、能量操控、战术分析 **5. 右下角区域(批次规格模块)** 包含一个金色警长星徽图案,内文:“活力 编号01 未来”。 - **大标题**(醒目文字):“ETB-77”。 - **描述文字**:“这是将分配给未来探险任务的批次。” - **详细信息框**(下方数据框):“批次规格|成员数量:2人|任务时长:365天|装备清单:完整|能源补给:充足|目标区域:未定”。 **6. 底部横幅与背景装饰(标语与元素模块)** 包含一段说明文字:“KATAO探险队|连接过去与未来|探索无限可能|官方网站:www.SENSETIME.com|联系我们:SENSETIME” 下方配有背景元素装饰:散落的黄色星星、星座线图案、技术网格线装饰。 整体视觉风格复古、结构清晰,通过鲜艳色彩与颗粒质感增强视觉冲击力和未来感,强调在未来探险任务中遵循精英组队、装备未来科技武装,并通过时空导航与能量操控达成探索无限可能的目标。所有文本均为中文(保留特定英文标识),语言简洁硬核,适合用于科幻探险主题的复古漫画风格海报宣传场景。" \
  --width 2720 --height 1536 \
  --cfg_scale 4.0 --cfg_norm none --timestep_shift 3.0 --num_steps 50 \
  --output output.png --profile

Default resolution is 2048×2048 (1:1). See supported resolution buckets for other aspect ratios.

For high-quality infographic generation, it is recommended to apply prompt enhancement before generating images.

✏️ Edit Infographics

python examples/editing/inference.py \
  --model_path sensenova/SenseNova-U1-8B-MoT-Infographic-V3 \
  --prompt "Change the main title to 'SenseNova-U1 V3' while preserving the remaining content and layout." \
  --image docs/assets/showcases/t2i_infographic/0004.webp \
  --cfg_scale 4.0 --img_cfg_scale 1.0 --cfg_norm none \
  --timestep_shift 3.0 --num_steps 50 \
  --output edited_infographic.png --profile --compare

💾 Memory-efficient inference (GGUF + VRAM modes)

For users running on a single consumer GPU, two complementary features lower the VRAM footprint of the transformers path. They can be combined freely.

--vram_mode: single-GPU layer offload

Pass --vram_mode to keep the language-model layers resident on CPU pinned memory and stream them onto the GPU on-demand during forward, freeing weight VRAM while keeping activations on-device.

Mode Behavior When to use
full (default) No offload; whole model on GPU Plenty of VRAM, best speed
low Synchronous per-layer CPU↔GPU swap Lowest VRAM footprint
balanced Async prefetch overlaps H2D copy with compute Tight on VRAM but want to recover speed
python examples/t2i/inference.py \
  --model_path sensenova/SenseNova-U1-8B-MoT-Infographic-V3 \
  --vram_mode balanced \
  --prompt "..." --output output.png

--gguf_checkpoint and --vram_mode compose: a Q4 GGUF + balanced is the recommended setup for ~10–12 GB consumer cards.

⚡ Run with LightLLM + LightX2V (Recommended)

For production serving, we co-design a dedicated inference stack on top of LightLLM (understanding) and LightX2V (generation). The two engines are disaggregated so that each path can use its own parallelism and resource budget, with a low-overhead transfer channel in between.

On a single node with TP2 + CFG2, this stack delivers roughly ~0.15 s/step and ~9 s end-to-end for a 2048×2048 image on H100 / H200, with a ~2.4–3.2× prefill speedup from our FA3-based hybrid-mask attention over the Triton baseline. Full per-GPU performance are reported in docs/inference_infra.md.

An official docker image is provided for one-command deployment:

docker pull lightx2v/lightllm_lightx2v:20260407

⚙️ Deployment guide (Docker, launch flags, modes, quantization, API test): see docs/deployment.md.

📖 Full design and performance profiling: see docs/inference_infra.md.

🌐 Join the Community!

Join our growing community to share feedback, get support, and stay updated on the latest SenseNova-U1 developments — we'd love to hear from you!

Discord WeChat Group

⚖️ License

This project is released under the Apache 2.0 License.

Downloads last month
365
Safetensors
Model size
18B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for sensenova/SenseNova-U1-8B-MoT-Infographic-V3

Finetunes
2 models

Space using sensenova/SenseNova-U1-8B-MoT-Infographic-V3 1

Paper for sensenova/SenseNova-U1-8B-MoT-Infographic-V3