pinecoresystems commited on
Commit
f9c7df9
·
verified ·
1 Parent(s): c47b925

Upload UPSTREAM-README.md with huggingface_hub

Browse files
Files changed (1) hide show
  1. UPSTREAM-README.md +94 -94
UPSTREAM-README.md CHANGED
@@ -1,94 +1,94 @@
1
- ---
2
- license: mit
3
- library_name: custom
4
- pipeline_tag: text-to-image
5
- inference: false
6
- tags:
7
- - text-to-image
8
- - image-generation
9
- - graphic-design
10
- - text-rendering
11
- - rgba
12
- ---
13
-
14
- # Ming-Image-0.1-Design
15
-
16
- [🧩 ModelScope](https://www.modelscope.cn/models/inclusionAI/Ming-Image-0.1-Design) · [🤗 Hugging Face](https://huggingface.co/inclusionAI/Ming-Image-0.1-Design) · [📄 Blog](https://mp.weixin.qq.com/s/VGdtxfM8kbHIQJw50VD_Sw) · [🖥️ Demo](https://huggingface.co/spaces/hugging-apps/ming-image-0-1-design-demo)<br>
17
- [🎨 Design Skill](https://github.com/inclusionAI/ling-cookbook/tree/main/resources/recommended-skills/ling-ui-design) · [📊 PPT Skill](https://github.com/inclusionAI/ling-cookbook/tree/main/resources/recommended-skills/image-to-editable-ppt)
18
-
19
- Ming-Image-0.1-Design is a 6B text-to-image model for UI, infographics,
20
- posters, and other text-rich visual designs. It generates complete visual
21
- compositions and supports RGBA output with transparent backgrounds.
22
-
23
- ## UI/UX Design leaderboard
24
-
25
- <p align="center">
26
- <img src="./assets/uiux_leaderboard.webp" width="100%" alt="Ming-Image-0.1-Design UI/UX Design leaderboard">
27
- </p>
28
-
29
- ## Quick Start
30
-
31
- Use the companion [Ming-Image repository](https://github.com/inclusionAI/Ming-Image)
32
- for installation and inference:
33
-
34
- ```bash
35
- git clone https://github.com/inclusionAI/Ming-Image
36
- cd Ming-Image
37
- pip install -r requirements.txt
38
-
39
- python infer.py \
40
- --model inclusionAI/Ming-Image-0.1-Design \
41
- --task text-to-image \
42
- --prompt assets/t2i_four_seasons_cabin_prompt.json \
43
- --resolution 2048 \
44
- --output-dir outputs/t2i
45
- ```
46
-
47
- Prompt enhancement (PE) can use `Ling-3.0-flash-VL` or `qwen3.8-27B`; see
48
- [text-to-image prompt rewriting](https://github.com/inclusionAI/Ming-Image#text-to-image-prompt-rewriting).
49
-
50
- ### Transparent-background generation
51
-
52
- For transparent-background generation, prepend exactly one of the recommended
53
- RGBA phrases. See the
54
- [transparent-background generation tip](https://github.com/inclusionAI/Ming-Image#transparent-background-generation-tip).
55
-
56
- ## Deployment
57
-
58
- We recommend the following inference frameworks to serve the model:
59
-
60
- - vLLM-Omni: see the [recipes](https://github.com/vllm-project/vllm-omni/blob/main/recipes/inclusionAI/Ming-Image.md)
61
- and [installation guide](https://docs.vllm.ai/projects/vllm-omni/en/latest/getting_started/quickstart/).
62
-
63
- ## Recommended settings
64
-
65
- - Resolution: **2048 x 2048** (recommended), or **1024 x 1024** for faster
66
- generation.
67
- - Sampling steps: **12**.
68
- - CFG scale: **1.0**.
69
- - Precision: **BF16**.
70
- - Hardware: **one CUDA GPU with 80 GiB VRAM** (validated configuration).
71
-
72
- The public inference code maps text-to-image resolution requests to the
73
- supported 1024 or 2048 bucket.
74
-
75
- ## Gallery
76
-
77
- ### Text-to-image
78
-
79
- <p align="center">
80
- <img src="./assets/showcase.webp" width="100%" alt="Ming-Image-0.1-Design generated examples">
81
- </p>
82
-
83
- ### Transparent-background text-to-image
84
-
85
- <p align="center">
86
- <img src="./assets/transparency_showcase.webp" width="100%" alt="Ming-Image-0.1-Design transparent-background examples">
87
- </p>
88
-
89
- The checkerboard is used only to preview transparency; it is not part of the
90
- generated RGBA images.
91
-
92
- ## License
93
-
94
- This model is released under the [MIT License](./LICENSE).
 
1
+ ---
2
+ license: mit
3
+ library_name: custom
4
+ pipeline_tag: text-to-image
5
+ inference: false
6
+ tags:
7
+ - text-to-image
8
+ - image-generation
9
+ - graphic-design
10
+ - text-rendering
11
+ - rgba
12
+ ---
13
+
14
+ # Ming-Image-0.1-Design
15
+
16
+ [🧩 ModelScope](https://www.modelscope.cn/models/inclusionAI/Ming-Image-0.1-Design) · [🤗 Hugging Face](https://huggingface.co/inclusionAI/Ming-Image-0.1-Design) · [📄 Blog](https://mp.weixin.qq.com/s/VGdtxfM8kbHIQJw50VD_Sw) · [🖥️ Demo](https://huggingface.co/spaces/hugging-apps/ming-image-0-1-design-demo)<br>
17
+ [🎨 Design Skill](https://github.com/inclusionAI/ling-cookbook/tree/main/resources/recommended-skills/ling-ui-design) · [📊 PPT Skill](https://github.com/inclusionAI/ling-cookbook/tree/main/resources/recommended-skills/image-to-editable-ppt)
18
+
19
+ Ming-Image-0.1-Design is a 6B text-to-image model for UI, infographics,
20
+ posters, and other text-rich visual designs. It generates complete visual
21
+ compositions and supports RGBA output with transparent backgrounds.
22
+
23
+ ## UI/UX Design leaderboard
24
+
25
+ <p align="center">
26
+ <img src="./assets/uiux_leaderboard.webp" width="100%" alt="Ming-Image-0.1-Design UI/UX Design leaderboard">
27
+ </p>
28
+
29
+ ## Quick Start
30
+
31
+ Use the companion [Ming-Image repository](https://github.com/inclusionAI/Ming-Image)
32
+ for installation and inference:
33
+
34
+ ```bash
35
+ git clone https://github.com/inclusionAI/Ming-Image
36
+ cd Ming-Image
37
+ pip install -r requirements.txt
38
+
39
+ python infer.py \
40
+ --model inclusionAI/Ming-Image-0.1-Design \
41
+ --task text-to-image \
42
+ --prompt assets/t2i_four_seasons_cabin_prompt.json \
43
+ --resolution 2048 \
44
+ --output-dir outputs/t2i
45
+ ```
46
+
47
+ Prompt enhancement (PE) can use `Ling-3.0-flash-VL` or `qwen3.8-27B`; see
48
+ [text-to-image prompt rewriting](https://github.com/inclusionAI/Ming-Image#text-to-image-prompt-rewriting).
49
+
50
+ ### Transparent-background generation
51
+
52
+ For transparent-background generation, prepend exactly one of the recommended
53
+ RGBA phrases. See the
54
+ [transparent-background generation tip](https://github.com/inclusionAI/Ming-Image#transparent-background-generation-tip).
55
+
56
+ ## Deployment
57
+
58
+ We recommend the following inference frameworks to serve the model:
59
+
60
+ - vLLM-Omni: see the [recipes](https://github.com/vllm-project/vllm-omni/blob/main/recipes/inclusionAI/Ming-Image.md)
61
+ and [installation guide](https://docs.vllm.ai/projects/vllm-omni/en/latest/getting_started/quickstart/).
62
+
63
+ ## Recommended settings
64
+
65
+ - Resolution: **2048 x 2048** (recommended), or **1024 x 1024** for faster
66
+ generation.
67
+ - Sampling steps: **12**.
68
+ - CFG scale: **1.0**.
69
+ - Precision: **BF16**.
70
+ - Hardware: **one CUDA GPU with 80 GiB VRAM** (validated configuration).
71
+
72
+ The public inference code maps text-to-image resolution requests to the
73
+ supported 1024 or 2048 bucket.
74
+
75
+ ## Gallery
76
+
77
+ ### Text-to-image
78
+
79
+ <p align="center">
80
+ <img src="./assets/showcase.webp" width="100%" alt="Ming-Image-0.1-Design generated examples">
81
+ </p>
82
+
83
+ ### Transparent-background text-to-image
84
+
85
+ <p align="center">
86
+ <img src="./assets/transparency_showcase.webp" width="100%" alt="Ming-Image-0.1-Design transparent-background examples">
87
+ </p>
88
+
89
+ The checkerboard is used only to preview transparency; it is not part of the
90
+ generated RGBA images.
91
+
92
+ ## License
93
+
94
+ This model is released under the [MIT License](./LICENSE).