SeeSee21 commited on
Commit
fdb7081
·
verified ·
1 Parent(s): 88fcae3

Restore original model card and extend it with AniSee V2

Browse files

The previous commit replaced the model card wholesale, dropping the badge header, cover image, V1 preview gallery, workflow section and notes. This restores all of them and folds the V2 content into the original structure instead.

Files changed (2) hide show
  1. README.md +387 -168
  2. config.json +177 -177
README.md CHANGED
@@ -20,79 +20,96 @@ tags:
20
  - anisee
21
  ---
22
 
 
 
23
  # 🎨 AniSee
24
 
25
- **A personal anime model built on Anima.**
26
- Two generations · Four checkpoints · Tags + natural language · ComfyUI-ready · LoRA-friendly
 
 
 
 
 
27
 
28
- AniSee started as a full fine-tune of *Anima Preview3 Base* and has since moved onto the final
29
- *Anima v1.0 / v1.1* releases. Both generations live in this repo, because they are genuinely
30
- different things and some people will prefer the older one.
 
 
 
 
 
 
 
 
 
31
 
32
- Everything you need is here — the models, the Qwen text encoder and the Qwen-Image VAE.
33
- No hunting across repos.
 
 
 
34
 
35
  ---
36
 
37
  ## ⬇️ Which file do I download?
38
 
39
- | File | Built on | Steps / CFG | Pick it if |
40
- |---|---|---|---|
41
- | **`AniSee-V2-Turbo.safetensors`** ⭐ | anima-turbo-v1.1 | 12 / 1 | You want speed. ~6 s per image. **Start here.** |
42
  | **`AniSee-V2-Aesthetic.safetensors`** | anima-aesthetic-v1.1 | 40 / 4–5 | You want the best look out of the box |
43
  | **`AniSee-V2-Base.safetensors`** | anima-base-v1.0 | 40 / 4–5 | You want maximum variety, or you train LoRAs |
44
  | `anisee.safetensors` | Anima Preview3 Base | 40 / 4.5 | You are already on the classic Preview3 setup |
45
 
46
- All four are **diffusion model files**. You also need the text encoder and the VAE from this
47
- repo — both are included below.
48
 
49
- > An **all-in-one** build that packs model, text encoder and VAE into a single checkpoint
50
- > exists for V1 on [CivitAI](https://civitai.com/models/2628747/anisee). A V2 all-in-one is
51
- > planned.
52
 
53
  ---
54
 
55
- ## 🖼️ Preview — the three V2 variants, same seed
56
-
57
- Each image below is one prompt rendered by all three V2 variants with the **seed locked
58
- identical**, so the only variable is the checkpoint. Order is **Base · Turbo · Aesthetic**,
59
- left to right.
60
 
61
- ![seesee_elf — aerial silk](images/v2/comparison/01-seesee-elf-aerial-silk.webp)
62
 
63
- ![seesee_kitsune — shrine](images/v2/comparison/03-seesee-kitsune-shrine.webp)
 
 
 
 
64
 
65
- ![Dragon woman — beach](images/v2/comparison/05-dragon-woman-beach.webp)
 
 
 
 
 
66
 
67
- ![Pure tag prompt — fox girl](images/v2/comparison/09-tags-fox-girl.webp)
68
-
69
- ![Pixel art — jungle](images/v2/comparison/04-pixel-art-jungle.webp)
70
 
71
- The remaining comparisons are in [`images/v2/comparison/`](images/v2/comparison), and each
72
- variant has its own 20-image gallery in [`images/v2/turbo/`](images/v2/turbo),
73
- [`images/v2/aesthetic/`](images/v2/aesthetic) and [`images/v2/base/`](images/v2/base).
74
 
75
- ---
 
 
76
 
77
- ## 🆕 What changed in V2
78
 
79
- V2 moves from Anima Preview3 onto the final Anima v1.0 / v1.1 releases, and it changes how the
80
- model was made.
81
 
82
- **V1 was a full fine-tune. V2 is not.** V2 is a **LoKr merge**: a LoKr network — linear 64 /
83
- alpha 64, conv 16 / alpha 16, full-rank, factor 4 — was trained for roughly 24,000 steps on
84
- `anima-base-v1.0`, then baked into three different Anima variants at strength 1.0. That is a
85
- lighter touch than a full fine-tune, and it is the honest description of what these files are.
86
 
87
- The upside of working that way: one training run gives three flavours. Same AniSee character,
88
- three foundations with genuinely different behaviour.
89
 
90
- ---
91
 
92
- ## 🔬 How the three variants actually differ
 
 
93
 
94
- The three were run side by side across ten prompts — seven natural language, three pure
95
- Danbooru tags — with the seed locked identical across all three models. What came out of it:
96
 
97
  - **Turbo diverges the most.** On the same seed it lands on a different pose and framing far
98
  more often than Base and Aesthetic differ from each other. That is the distillation of the
@@ -107,116 +124,184 @@ Danbooru tags — with the seed locked identical across all three models. What c
107
 
108
  ---
109
 
110
- ## 🔧 Installation (ComfyUI)
111
 
112
- Download the checkpoint you want plus the text encoder and the VAE, and drop them here:
 
113
 
114
- ```
115
- ComfyUI/models/diffusion_models/ AniSee-V2-Turbo.safetensors
116
- ComfyUI/models/text_encoders/ qwen_3_06b_base.safetensors
117
- ComfyUI/models/vae/ qwen_image_vae.safetensors
118
- ```
119
 
120
- Then wire up the standard Anima workflow:
 
 
121
 
122
- - **Load Diffusion Model** → `AniSee-V2-Turbo.safetensors`
123
- - **Load CLIP** → `qwen_3_06b_base.safetensors` (type: `stable_diffusion`)
124
- - **Load VAE** → `qwen_image_vae.safetensors`
 
 
 
125
 
126
- ⚠️ These are diffusion model files, so use **Load Diffusion Model**, *not* the Checkpoint
127
- Loader. If you already run Anima you have the text encoder and the VAE already.
128
 
129
- ### With `huggingface_hub`
 
 
 
 
 
130
 
131
- ```python
132
- from huggingface_hub import hf_hub_download
133
 
134
- model = hf_hub_download("SeeSee21/AniSee", "AniSee-V2-Turbo.safetensors")
135
- te = hf_hub_download("SeeSee21/AniSee", "text_encoders/qwen_3_06b_base.safetensors")
136
- vae = hf_hub_download("SeeSee21/AniSee", "vae/qwen_image_vae.safetensors")
137
- ```
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
138
 
139
  ---
140
 
141
- ## ⚙️ Recommended settings
142
 
143
- Sampler `er_sde` with scheduler `simple` for every variant — neutral style, flat colours,
144
- sharp lines.
145
 
146
  | Checkpoint | Steps | CFG | Note |
147
- |---|---|---|---|
148
- | AniSee-V2-Turbo | 12 | 1 | Negative prompt has no effect at CFG 1 |
149
- | AniSee-V2-Aesthetic | 40 | 4–5 | Leave out `score_*` tags |
150
- | AniSee-V2-Base | 40 | 4–5 | Most neutral, most flexible |
151
- | anisee (v1) | 40 | 4.5 | The original recommendation |
 
 
 
 
 
 
 
 
 
152
 
153
- **CFG guide.** 4.0–5.0 is the sweet spot for the 40-step variants. Above 5.0 you start risking
154
- burnt images, especially with heavy quality tags. If results feel harsh, drop CFG slightly or
155
- cut down on quality tags.
156
 
157
  **On Aesthetic:** per the Anima documentation, skip `score_*` tags entirely — in both the
158
  positive and the negative prompt. The base is already high quality and score tags push it into
159
  slop territory. `masterpiece, best quality` is fine to keep.
160
 
161
- ### Sampler alternatives
162
 
163
- | Sampler | Character |
164
- |---|---|
165
- | `er_sde` + `simple` | The default. Neutral, flat colours, sharp lines |
166
- | `euler_a` | Softer, thinner lines, slightly 2.5D, tolerates higher CFG |
167
- | `dpmpp_2m_sde_gpu` | Similar to er_sde but more creative; can get wild on short prompts |
168
  | `euler` | A bit more creative. Good on Turbo and Aesthetic, which are naturally stable |
169
 
 
 
170
  ---
171
 
172
- ## 📐 Resolution
173
 
174
- | Use case | Resolution |
175
- |---|---|
176
- | ⭐ Square / general purpose | 1024 × 1024 |
177
- | Portrait / character art | 896 × 1152 or 832 × 1216 |
178
- | Landscape / scenes | 1152 × 896 |
179
- | Wider cinematic | 1254 × 836 |
180
- | Widescreen | 1365 × 768 |
181
 
182
- Stay around 1 MP for the cleanest results. Anima works between 512² and 1536², but starts
183
- breaking down somewhere around 2 MP — generate at 1 MP and upscale afterwards for bigger images.
184
 
185
  ---
186
 
187
- ## 💡 Prompting
 
 
 
 
 
 
188
 
189
- AniSee inherits Anima's prompting system and accepts Danbooru-style tags, natural language, and
190
- any mix of the two.
191
 
192
  ```
193
- [quality tags] [meta tags] [safety tag] [subject] [character]
194
- [appearance] [pose] [clothing] [background] [lighting] [style]
195
  ```
196
 
197
- ### Tag rules inherited from Anima
198
 
199
- - Lowercase tags, spaces instead of underscores
200
- - Score tags are the only ones using underscores, e.g. `score_7`
201
- - Artist tags need an `@` prefix, e.g. `@artistname` — without it the effect is very weak
 
202
  - Where a tag differs between Danbooru and Gelbooru, prefer the Gelbooru spelling
203
  - Prompt weighting works but needs heavier weights than SDXL, e.g. `(chibi:2)`
204
 
205
- ### 🔑 V2 trigger words
206
 
207
- The V2 dataset was captioned in **natural language only**, so these want to be used in
208
- sentences rather than dropped in as bare tags:
209
 
210
  | Trigger | Concept |
211
- |---|---|
212
  | `seesee_elf` | White-haired elf |
213
  | `seesee_kitsune` | Fox woman |
214
 
215
- Give them a described setting, two sentences minimum. Very short prompts produce unpredictable
216
- results on Anima — the model fills the gaps with its own biases, and you may not like what it
217
- picks.
218
 
219
- ### ✅ Good — mixed prompt
220
 
221
  ```
222
  masterpiece, best quality, score_7, highres, illustration, safe, 1girl,
@@ -225,7 +310,7 @@ at night, neon lights reflecting on wet asphalt, cinematic lighting,
225
  detailed anime illustration
226
  ```
227
 
228
- ### ✅ Good — natural language
229
 
230
  ```
231
  masterpiece, best quality, score_7, highres, illustration.
@@ -238,21 +323,43 @@ detailed fabric shading, calm serene expression.
238
 
239
  ### ❌ Avoid
240
 
 
 
241
  ```
242
  anime girl, silver hair, hoodie
243
  ```
244
 
245
- Too sparse. Aim for a decent set of descriptive tags, or two or more sentences.
 
 
 
 
 
 
246
 
247
- ### ⭐ Recommended positive prefix
248
 
249
  ```
250
  masterpiece, best quality, score_7, highres, illustration,
251
  ```
252
 
253
- On **Aesthetic**, drop the `score_7` and use `masterpiece, best quality,` on its own.
 
 
 
 
 
 
 
 
 
 
 
 
 
254
 
255
- ### ⭐ Recommended negative prompt
 
256
 
257
  ```
258
  worst quality, low quality, score_1, score_2, score_3, artist name,
@@ -262,102 +369,214 @@ patreon username, web address, signature, watermark, artist name,
262
  censored, mosaic censoring
263
  ```
264
 
265
- If images come out flat or lose style, ease off the heavy weights — drop `(low quality:1.4)`
266
- back to plain `low quality`. On Aesthetic, remove the `score_*` terms. On Turbo at CFG 1 it is
267
- ignored entirely.
 
 
 
 
268
 
269
- ### 🛡️ Safety tags
270
 
271
- Inherited from Anima. Use one in the positive prompt: `safe` (recommended default),
272
- `sensitive`, `nsfw`, or `explicit`.
 
 
273
 
274
  ---
275
 
276
- ## 📁 Repository structure
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
277
 
 
 
278
  ```
279
- AniSee-V2-Turbo.safetensors 3.9 GB V2, distilled foundation, 12 steps
280
- AniSee-V2-Aesthetic.safetensors 3.9 GB V2, aesthetic foundation, 40 steps
281
- AniSee-V2-Base.safetensors 3.9 GB V2, neutral foundation, 40 steps
282
- anisee.safetensors 3.9 GB V1, full fine-tune of Preview3
283
-
284
- text_encoders/qwen_3_06b_base.safetensors Qwen3-0.6B text encoder
285
- vae/qwen_image_vae.safetensors Qwen-Image VAE
286
-
287
- images/v2/comparison/ 10 three-way comparisons
288
- images/v2/turbo/ 20 samples
289
- images/v2/aesthetic/ 20 samples
290
- images/v2/base/ 20 samples
291
- images/ V1 gallery
292
-
293
- workflows/ ComfyUI workflows
294
- config.json Model metadata
 
 
 
 
 
 
 
 
 
 
295
  ```
296
 
297
  ---
298
 
299
- ## 📈 Version history
300
 
301
- ### V2 — LoKr merge onto Anima v1.0 / v1.1
 
 
302
 
303
- - Three variants released together: Turbo, Aesthetic, Base
304
- - LoKr network — linear 64 / alpha 64, conv 16 / alpha 16, full-rank, factor 4
305
- - ~24,000 training steps on `anima-base-v1.0`, merged at strength 1.0
306
- - Natural-language-only dataset; trigger words `seesee_elf` and `seesee_kitsune`
307
- - Diffusion model files — text encoder and VAE included in this repo
308
- - Moves off Anima Preview3 onto the final Anima v1.0 / v1.1 releases
309
 
310
- ### v1.0 — initial release
311
 
312
- - Full fine-tune of Anima Preview3 Base, ~20,000 steps on a curated anime dataset
313
- - LLM adapter only very lightly co-trained, per Anima's fine-tuning guidelines
314
- - Drop-in replacement for `anima-preview3-base.safetensors`
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
315
 
316
  ---
317
 
318
- ## 🗺️ Roadmap
 
 
319
 
320
- **✅ Released** — AniSee v1, AniSee V2 (Turbo, Aesthetic, Base)
 
 
 
 
 
321
 
322
- **🔜 Planned**
323
 
324
- - **AniSee V2 AIO** — an all-in-one build of V2, so V2 gets the same one-file convenience the
325
- V1 AIO has on CivitAI
326
- - **Official AniSee ComfyUI workflow** — a dedicated workflow, though the standard Anima
327
- workflow already covers both generations
 
 
 
328
 
329
  ---
330
 
331
  ## 🔗 Links
332
 
333
- - [AniSee on CivitAI](https://civitai.com/models/2628747/anisee) — including the V1 all-in-one build
334
- - [AniSee sample gallery](https://anisee.anisee.workers.dev/)
335
- - [Anima by CircleStone Labs](https://huggingface.co/circlestone-labs/Anima) — the base model
 
336
 
337
  ---
338
 
339
  ## 🙏 Credits
340
 
341
- - **Base model:** Anima by CircleStone Labs and Comfy Org — Preview3 Base for V1, v1.0 / v1.1 for V2
342
- - **Architecture:** built on NVIDIA Cosmos-Predict2-2B; Anima is a derivative model
343
- - **Fine-tune and merges:** SeeSee21
 
 
344
 
345
  ---
346
 
347
  ## 📜 License
348
 
349
- AniSee inherits the **CircleStone Labs Non-Commercial License** from Anima. The model and its
350
- derivatives may be used for non-commercial purposes only. As a derivative of
351
- Cosmos-Predict2-2B-Text2Image, the **NVIDIA Open Model License Agreement** also applies insofar
352
- as it covers derivative models.
353
 
354
- **Generated images are not covered by that restriction.** You may use the images you make
355
- commercially — selling images, paid commissions, concept art or assets for a paid product are
356
  all fine. What needs a separate license is hosting the model behind a paid API, embedding the
357
  weights in a monetized product, or running it on a paid generation platform.
358
 
359
- For commercial licensing of the base model, contact CircleStone Labs at `tdrussell@circlestone.ai`.
 
360
 
361
  ---
362
 
363
- *AniSee — a personal anime model built on Anima.* 🎨
 
 
 
 
 
 
 
 
 
 
20
  - anisee
21
  ---
22
 
23
+ <div align="center">
24
+
25
  # 🎨 AniSee
26
 
27
+ ### Personal Anime Model built on Anima
28
+
29
+ **Two Generations • Clean Anime Aesthetics • Tag + Natural Language • Anima-Compatible**
30
+
31
+ **Diffusion Models • 1 MP Native • LoRA-friendly**
32
+
33
+ <br>
34
 
35
+ <a href="https://civitai.red/models/2628747/anisee">
36
+ <img src="https://img.shields.io/badge/CivitAI-AniSee-EC4899?style=for-the-badge&logoColor=white" alt="CivitAI">
37
+ </a>
38
+ <a href="https://anisee.anisee.workers.dev/">
39
+ <img src="https://img.shields.io/badge/🎨_Sample_Gallery-Explore-D63384?style=for-the-badge" alt="Sample Gallery">
40
+ </a>
41
+ <a href="https://huggingface.co/circlestone-labs/Anima">
42
+ <img src="https://img.shields.io/badge/Base_Model-Anima-FFD21E?style=for-the-badge&logo=huggingface&logoColor=black" alt="Base Model: Anima">
43
+ </a>
44
+ <a href="#-license">
45
+ <img src="https://img.shields.io/badge/License-Non--Commercial-A02060?style=for-the-badge" alt="License">
46
+ </a>
47
 
48
+ <br><br>
49
+
50
+ <img src="https://huggingface.co/SeeSee21/AniSee/resolve/main/images/cover.png" alt="AniSee Cover" width="100%">
51
+
52
+ </div>
53
 
54
  ---
55
 
56
  ## ⬇️ Which file do I download?
57
 
58
+ | Checkpoint | Built on | Steps / CFG | Pick it if |
59
+ | --- | --- | --- | --- |
60
+ | **`AniSee-V2-Turbo.safetensors`** ⭐ | anima-turbo-v1.1 | **12 / 1** | You want speed — ~6 s per image. **Start here.** |
61
  | **`AniSee-V2-Aesthetic.safetensors`** | anima-aesthetic-v1.1 | 40 / 4–5 | You want the best look out of the box |
62
  | **`AniSee-V2-Base.safetensors`** | anima-base-v1.0 | 40 / 4–5 | You want maximum variety, or you train LoRAs |
63
  | `anisee.safetensors` | Anima Preview3 Base | 40 / 4.5 | You are already on the classic Preview3 setup |
64
 
65
+ All four are **diffusion model files**. The **Qwen text encoder** and the **Qwen-Image VAE** are
66
+ included in this repo — no hunting across repos.
67
 
68
+ > An **all-in-one** build that packs model, text encoder and VAE into a single checkpoint exists
69
+ > for V1 on [CivitAI](https://civitai.red/models/2628747/anisee). A V2 all-in-one is planned.
 
70
 
71
  ---
72
 
73
+ ## 🖼️ Preview Gallery
 
 
 
 
74
 
75
+ Browse the full curated set of sample images on the dedicated gallery page:
76
 
77
+ <div align="center">
78
+ <a href="https://anisee.anisee.workers.dev/">
79
+ <img src="https://img.shields.io/badge/🎨_Open_Sample_Gallery-anisee.anisee.workers.dev-D63384?style=for-the-badge" alt="Sample Gallery">
80
+ </a>
81
+ </div>
82
 
83
+ | | | |
84
+ | :---: | :---: | :---: |
85
+ | ![AniSee preview 1](https://huggingface.co/SeeSee21/AniSee/resolve/main/images/1.png) | ![AniSee preview 2](https://huggingface.co/SeeSee21/AniSee/resolve/main/images/2.png) | ![AniSee preview 3](https://huggingface.co/SeeSee21/AniSee/resolve/main/images/3.png) |
86
+ | ![AniSee preview 4](https://huggingface.co/SeeSee21/AniSee/resolve/main/images/4.png) | ![AniSee preview 5](https://huggingface.co/SeeSee21/AniSee/resolve/main/images/5.png) | ![AniSee preview 6](https://huggingface.co/SeeSee21/AniSee/resolve/main/images/6.webp) |
87
+ | ![AniSee preview 7](https://huggingface.co/SeeSee21/AniSee/resolve/main/images/7.webp) | ![AniSee preview 8](https://huggingface.co/SeeSee21/AniSee/resolve/main/images/8.png) | ![AniSee preview 9](https://huggingface.co/SeeSee21/AniSee/resolve/main/images/9.png) |
88
+ | ![AniSee preview 10](https://huggingface.co/SeeSee21/AniSee/resolve/main/images/10.png) | ![AniSee preview 11](https://huggingface.co/SeeSee21/AniSee/resolve/main/images/11.webp) | ![AniSee preview 12](https://huggingface.co/SeeSee21/AniSee/resolve/main/images/12.webp) |
89
 
90
+ ---
 
 
91
 
92
+ ## 🔬 V2 Side by Side — Base · Turbo · Aesthetic
 
 
93
 
94
+ Each image below is **one prompt rendered by all three V2 variants with the seed locked
95
+ identical**, so the only variable is the checkpoint. Order is **Base · Turbo · Aesthetic**,
96
+ left to right.
97
 
98
+ <img src="https://huggingface.co/SeeSee21/AniSee/resolve/main/images/v2/comparison/01-seesee-elf-aerial-silk.webp" alt="seesee_elf — aerial silk" width="100%">
99
 
100
+ <img src="https://huggingface.co/SeeSee21/AniSee/resolve/main/images/v2/comparison/03-seesee-kitsune-shrine.webp" alt="seesee_kitsune — shrine" width="100%">
 
101
 
102
+ <img src="https://huggingface.co/SeeSee21/AniSee/resolve/main/images/v2/comparison/05-dragon-woman-beach.webp" alt="Dragon woman — beach" width="100%">
 
 
 
103
 
104
+ <img src="https://huggingface.co/SeeSee21/AniSee/resolve/main/images/v2/comparison/09-tags-fox-girl.webp" alt="Pure tag prompt — fox girl" width="100%">
 
105
 
106
+ <img src="https://huggingface.co/SeeSee21/AniSee/resolve/main/images/v2/comparison/04-pixel-art-jungle.webp" alt="Pixel art — jungle" width="100%">
107
 
108
+ The remaining comparisons live in [`images/v2/comparison/`](./images/v2/comparison), and each
109
+ variant has its own 20-image gallery in [`images/v2/turbo/`](./images/v2/turbo),
110
+ [`images/v2/aesthetic/`](./images/v2/aesthetic) and [`images/v2/base/`](./images/v2/base).
111
 
112
+ **What the ten test prompts showed:**
 
113
 
114
  - **Turbo diverges the most.** On the same seed it lands on a different pose and framing far
115
  more often than Base and Aesthetic differ from each other. That is the distillation of the
 
124
 
125
  ---
126
 
127
+ ## ✨ What is AniSee?
128
 
129
+ AniSee is a personal anime model built on **Anima** by CircleStone Labs, retrained on my own
130
+ curated dataset to push the model further into a cleaner, more focused anime aesthetic.
131
 
132
+ It exists in two generations, and both live in this repo:
 
 
 
 
133
 
134
+ **V1** is a **full fine-tune** of *Anima Preview3 Base* — around 20K training steps, with the
135
+ LLM adapter only very lightly co-trained, following the official Anima fine-tuning guidelines.
136
+ It is not a LoRA merge.
137
 
138
+ **V2** moves onto the final *Anima v1.0 / v1.1* releases and uses a different method. V2 is a
139
+ **LoKr merge**: a LoKr network — linear 64 / alpha 64, conv 16 / alpha 16, full-rank, factor 4 —
140
+ was trained for roughly 24,000 steps on `anima-base-v1.0`, then baked into three different Anima
141
+ variants at strength 1.0. That is a lighter touch than a full fine-tune, and it is the honest
142
+ description of what those files are. The upside: one training run gives three flavours, each
143
+ with genuinely different behaviour.
144
 
145
+ The goal in both generations is to keep everything that makes Anima a strong illustration base:
 
146
 
147
+ - Danbooru-style tags
148
+ - Natural language prompts
149
+ - Mixed prompts
150
+ - Full Qwen text encoder support
151
+ - Qwen-Image VAE
152
+ - Anima-compatible generation behavior
153
 
154
+ while shifting the default style toward a stronger, cleaner anime look in line with my other
155
+ checkpoints.
156
 
157
+ AniSee is mainly intended for:
158
+
159
+ - Anime-style illustrations
160
+ - Character-focused images
161
+ - Cleaner anime aesthetics
162
+ - Style experiments
163
+ - Testing Anima-based fine-tunes inside ComfyUI
164
+
165
+ ---
166
+
167
+ ## 🎯 Key Features
168
+
169
+ - ✅ **Three V2 variants** from one training run — Turbo, Aesthetic and Base
170
+ - ✅ **Turbo runs at 12 steps / CFG 1** — around 6 seconds per image
171
+ - ✅ V1 remains available as a **full fine-tune** of Anima Preview3 Base
172
+ - ✅ Clean, focused anime aesthetics
173
+ - ✅ Supports Danbooru-style tags, natural language, and mixed prompts
174
+ - ✅ Compatible with the standard Anima ComfyUI workflow
175
+ - ✅ Uses the existing **Qwen text encoder** + **Qwen-Image VAE** — included in the repo
176
+ - ✅ LoRA training friendly — same base architecture as Anima
177
+ - ✅ Official ComfyUI workflow included, with auto quality prefix and Qwen3-VL prompt enhancer
178
+
179
+ ---
180
+
181
+ ## 🗺️ AniSee Roadmap
182
+
183
+ ### ✅ Released
184
+
185
+ #### 🎨 AniSee V1
186
+
187
+ Full fine-tune of Anima Preview3 Base — Diffusion Model variant. The foundation of the AniSee
188
+ family. An all-in-one build of V1 is available on CivitAI.
189
+
190
+ #### 🚀 AniSee V2 — Turbo, Aesthetic and Base
191
+
192
+ LoKr merge onto the final Anima v1.0 / v1.1 releases, in three variants.
193
+
194
+ #### 🔧 Official AniSee ComfyUI Workflow
195
+
196
+ A dedicated workflow with the auto-prefix, optional Qwen3-VL prompt enhancer, LoRA support and
197
+ Ultimate SD Upscale is included in this repo under
198
+ [`workflows/anisee-workflow-SDUltimate.json`](./workflows/anisee-workflow-SDUltimate.json).
199
+
200
+ ### 🔜 Planned
201
+
202
+ #### 📦 AniSee V2 AIO
203
+
204
+ All-in-one V2 checkpoint with **Diffusion Model + Qwen Text Encoder + Qwen-Image VAE** in a
205
+ single file, so V2 gets the same one-file convenience the V1 AIO already has.
206
+
207
+ More updates coming as testing progresses! 🎨
208
 
209
  ---
210
 
211
+ ## ⚙️ Recommended Settings
212
 
213
+ Sampler `er_sde` with scheduler `simple` for every variant — neutral style, flat colors, sharp
214
+ lines.
215
 
216
  | Checkpoint | Steps | CFG | Note |
217
+ | --- | :---: | :---: | --- |
218
+ | **AniSee-V2-Turbo** | 12 | 1 | The negative prompt has no effect at CFG 1 |
219
+ | **AniSee-V2-Aesthetic** | 40 | 4–5 | Leave out `score_*` tags |
220
+ | **AniSee-V2-Base** | 40 | 4–5 | Most neutral, most flexible |
221
+ | `anisee` (V1) | 40 | 4.5 | The original recommendation |
222
+
223
+ ```yaml
224
+ # AniSee-V2-Turbo
225
+ Steps: 12
226
+ CFG: 1
227
+ Sampler: er_sde
228
+ Scheduler: simple
229
+ Resolution: ~1 MP # e.g. 832×1216, 896×1152, 1024×1024
230
+ ```
231
 
232
+ **CFG Guide:** 4.0–5.0 is the sweet spot for the 40-step variants. Going above 5.0 starts to
233
+ risk burning the image, especially with heavy quality tags. If results feel too harsh, drop CFG
234
+ slightly or reduce quality tag count.
235
 
236
  **On Aesthetic:** per the Anima documentation, skip `score_*` tags entirely — in both the
237
  positive and the negative prompt. The base is already high quality and score tags push it into
238
  slop territory. `masterpiece, best quality` is fine to keep.
239
 
240
+ **Sampler alternatives** (all work well, just different character):
241
 
242
+ | Sampler / Scheduler | Character |
243
+ | --- | --- |
244
+ | `er_sde` + `simple` *(default)* | Neutral style, flat colors, sharp lines |
245
+ | `euler_a` | Softer, thinner lines, slightly more 2.5D feel, tolerates higher CFG |
246
+ | `dpmpp_2m_sde_gpu` | Similar to er_sde but more "creative", can get wild on short prompts |
247
  | `euler` | A bit more creative. Good on Turbo and Aesthetic, which are naturally stable |
248
 
249
+ Feel free to experiment — these are just starting points, not hard rules.
250
+
251
  ---
252
 
253
+ ## 📐 Resolution Guide
254
 
255
+ | Use Case | Resolution |
256
+ | --- | --- |
257
+ | ⭐ Square / General purpose | **1024 × 1024** |
258
+ | Portrait / Character art | **896 × 1152** or **832 × 1216** |
259
+ | Landscape / Scenes | **1152 × 896** |
260
+ | Wider cinematic | **1254 × 836** |
261
+ | Widescreen | **1365 × 768** |
262
 
263
+ Stay around **1 MP** for the cleanest results. The Anima base starts breaking down somewhere
264
+ around 2 MP, so if you want bigger images, generate at 1 MP first and upscale afterwards.
265
 
266
  ---
267
 
268
+ ## 💡 Prompting Guide
269
+
270
+ AniSee inherits Anima's prompting system. It accepts:
271
+
272
+ - Danbooru / anime-style tags
273
+ - Natural language prompts
274
+ - Mixed prompts with tags + sentences
275
 
276
+ A good prompt structure:
 
277
 
278
  ```
279
+ [quality tags] [meta tags] [safety tag] [subject] [character] [appearance]
280
+ [pose] [clothing] [background] [lighting] [style]
281
  ```
282
 
283
+ **Important tag rules inherited from Anima**
284
 
285
+ - Use lowercase for tags, spaces instead of underscores
286
+ - Score tags are the only tags that use underscores, for example `score_7`
287
+ - Artist tags must be prefixed with `@`, for example `@artistname` — without it the effect is
288
+ very weak
289
  - Where a tag differs between Danbooru and Gelbooru, prefer the Gelbooru spelling
290
  - Prompt weighting works but needs heavier weights than SDXL, e.g. `(chibi:2)`
291
 
292
+ ### 🔑 V2 Trigger Words
293
 
294
+ The V2 dataset was captioned in **natural language only**, so these want to be used inside
295
+ descriptive sentences rather than dropped in as bare tags:
296
 
297
  | Trigger | Concept |
298
+ | --- | --- |
299
  | `seesee_elf` | White-haired elf |
300
  | `seesee_kitsune` | Fox woman |
301
 
302
+ Give them a described setting, two sentences minimum.
 
 
303
 
304
+ ### ✅ Good (mixed prompt)
305
 
306
  ```
307
  masterpiece, best quality, score_7, highres, illustration, safe, 1girl,
 
310
  detailed anime illustration
311
  ```
312
 
313
+ ### ✅ Good (natural language)
314
 
315
  ```
316
  masterpiece, best quality, score_7, highres, illustration.
 
323
 
324
  ### ❌ Avoid
325
 
326
+ Very short tag dumps like:
327
+
328
  ```
329
  anime girl, silver hair, hoodie
330
  ```
331
 
332
+ The model can produce unexpected results when the prompt is too sparse — it fills the gaps with
333
+ its own biases, and you may not like what it picks. Aim for at least a few descriptive tags or
334
+ 2+ sentences.
335
+
336
+ ---
337
+
338
+ ## ⭐ Recommended Positive Prefix
339
 
340
+ Start every prompt with:
341
 
342
  ```
343
  masterpiece, best quality, score_7, highres, illustration,
344
  ```
345
 
346
+ Then add your subject, character, scene, and style tags after that. On **AniSee-V2-Aesthetic**,
347
+ drop the `score_7` and use `masterpiece, best quality,` on its own.
348
+
349
+ You can also experiment with other quality tag combinations:
350
+
351
+ ```
352
+ masterpiece, best quality, score_7, safe
353
+ masterpiece, best quality, score_8, highres, official art
354
+ score_9, masterpiece, absurdres, anime screenshot
355
+ ```
356
+
357
+ ---
358
+
359
+ ## ⭐ Recommended Negative Prompt
360
 
361
+ This is the negative prompt I run with — it cleans up most common issues without being so
362
+ aggressive that it kills the style:
363
 
364
  ```
365
  worst quality, low quality, score_1, score_2, score_3, artist name,
 
369
  censored, mosaic censoring
370
  ```
371
 
372
+ If your images come out too flat or lose style, reduce the weights on the heavier terms, for
373
+ example drop `(low quality:1.4)` back to `low quality`. On Aesthetic, remove the `score_*`
374
+ terms. On Turbo at CFG 1 the negative prompt is ignored entirely.
375
+
376
+ ---
377
+
378
+ ## 🛡️ Safety Tags
379
 
380
+ Inherited from Anima. Use one of these in the positive prompt:
381
 
382
+ - `safe` — for normal generations, recommended default
383
+ - `sensitive`
384
+ - `nsfw`
385
+ - `explicit`
386
 
387
  ---
388
 
389
+ ## 🔧 Installation
390
+
391
+ ### Step 1 — Download the files
392
+
393
+ You need three files (all included in this repo):
394
+
395
+ - One checkpoint — e.g. `AniSee-V2-Turbo.safetensors`
396
+ - `text_encoders/qwen_3_06b_base.safetensors` — text encoder
397
+ - `vae/qwen_image_vae.safetensors` — VAE
398
+
399
+ ### Step 2 — Place the files
400
+
401
+ ```
402
+ ComfyUI/models/diffusion_models/
403
+ └── AniSee-V2-Turbo.safetensors
404
+
405
+ ComfyUI/models/text_encoders/
406
+ └── qwen_3_06b_base.safetensors
407
 
408
+ ComfyUI/models/vae/
409
+ └── qwen_image_vae.safetensors
410
  ```
411
+
412
+ If you already run **Anima**, you already have the text encoder and VAE — AniSee is a direct
413
+ drop-in.
414
+
415
+ ### Step 3 — Load in ComfyUI
416
+
417
+ Use the standard Anima workflow, or the official AniSee workflow from
418
+ `workflows/anisee-workflow-SDUltimate.json`:
419
+
420
+ - **Load Diffusion Model** → `AniSee-V2-Turbo.safetensors`
421
+ - **Load Text Encoder** → `qwen_3_06b_base.safetensors`
422
+ - **Load VAE** → `qwen_image_vae.safetensors`
423
+
424
+ Then your usual sampler, encode, decode, save chain.
425
+
426
+ ⚠️ These are **diffusion model** files, so use **Load Diffusion Model** — *not* the Checkpoint
427
+ Loader.
428
+
429
+ ### With `huggingface_hub`
430
+
431
+ ```python
432
+ from huggingface_hub import hf_hub_download
433
+
434
+ model = hf_hub_download("SeeSee21/AniSee", "AniSee-V2-Turbo.safetensors")
435
+ te = hf_hub_download("SeeSee21/AniSee", "text_encoders/qwen_3_06b_base.safetensors")
436
+ vae = hf_hub_download("SeeSee21/AniSee", "vae/qwen_image_vae.safetensors")
437
  ```
438
 
439
  ---
440
 
441
+ ## 🧩 Official Workflow
442
 
443
+ <div align="center">
444
+ <img src="https://huggingface.co/SeeSee21/AniSee/resolve/main/images/anisee-workflow-cover.png" alt="AniSee Workflow" width="100%">
445
+ </div>
446
 
447
+ A ready-to-use ComfyUI workflow is included at [`workflows/anisee-workflow-SDUltimate.json`](./workflows/anisee-workflow-SDUltimate.json).
 
 
 
 
 
448
 
449
+ It features:
450
 
451
+ - 📦 Model + Text Encoder + VAE loaders pre-configured
452
+ - 🔗 **Auto Quality Prefix** — no need to type `masterpiece, best quality, score_7, ...` yourself
453
+ - 🎲 **Optional Qwen3-VL Prompt Enhancer** — converts short one-liners into full Danbooru tag lists
454
+ - 📖 Optional **LoRA** stack via Lora Manager (one-click toggle)
455
+ - 🔼 Optional **UltimateSDUpscale** 2× with side-by-side compare
456
+ - 🎨 Pre-configured with `er_sde` / `simple` / 40 steps / CFG 4.5
457
+ - ➖ Pre-loaded recommended negative prompt
458
+ - 📝 Built-in MarkdownNote with all settings + quick reference
459
+
460
+ > Using **AniSee-V2-Turbo**? Set the sampler to **12 steps / CFG 1** — the workflow ships with
461
+ > the 40-step / CFG 4.5 defaults for the Base and Aesthetic variants.
462
+
463
+ **Required custom nodes** (all installable via ComfyUI Manager):
464
+
465
+ - [ComfyUI-Easy-Use](https://github.com/yolain/ComfyUI-Easy-Use)
466
+ - [ComfyUI_UltimateSDUpscale](https://github.com/ssitu/ComfyUI_UltimateSDUpscale)
467
+ - [ComfyUI-Lora-Manager](https://github.com/willmiao/ComfyUI-Lora-Manager)
468
+ - [ComfyUI-QwenVL](https://github.com/1038lab/ComfyUI-QwenVL)
469
+ - [rgthree-comfy](https://github.com/rgthree/rgthree-comfy)
470
+
471
+ For the optional 2× upscaler, also place `4x-UltraSharp.pth` in `ComfyUI/models/upscale_models/`:
472
+
473
+ - [OpenModelDB — 4x-UltraSharp](https://openmodeldb.info/models/4x-UltraSharp)
474
+ - [HuggingFace — Kim2091/UltraSharp](https://huggingface.co/Kim2091/UltraSharp)
475
+
476
+ ---
477
+
478
+ ## 📁 Repository Structure
479
+
480
+ ```
481
+ AniSee/
482
+ ├── README.md
483
+ ├── config.json
484
+ │
485
+ ├── AniSee-V2-Turbo.safetensors # V2, distilled foundation, 12 steps (~3.9 GB)
486
+ ├── AniSee-V2-Aesthetic.safetensors # V2, aesthetic foundation, 40 steps (~3.9 GB)
487
+ ├── AniSee-V2-Base.safetensors # V2, neutral foundation, 40 steps (~3.9 GB)
488
+ ├── anisee.safetensors # V1, full fine-tune of Preview3 (~3.9 GB)
489
+ │
490
+ ├── text_encoders/
491
+ │ └── qwen_3_06b_base.safetensors # text encoder (same as Anima)
492
+ │
493
+ ├── vae/
494
+ │ └── qwen_image_vae.safetensors # VAE (same as Anima)
495
+ │
496
+ ├── images/
497
+ │ ├── cover.png # social preview / model cover
498
+ │ ├── anisee-workflow-cover.png # workflow preview image
499
+ │ ├── 1.png 2.png 3.png 4.png # V1 gallery
500
+ │ ├── 5.png 6.webp 7.webp 8.png
501
+ │ ├── 9.png 10.png 11.webp 12.webp
502
+ │ └── v2/
503
+ │ ├── comparison/ # 10 three-way comparisons
504
+ │ ├── turbo/ # 20 samples
505
+ │ ├── aesthetic/ # 20 samples
506
+ │ └── base/ # 20 samples
507
+ │
508
+ └── workflows/
509
+ └── anisee-workflow-SDUltimate.json
510
+ ```
511
 
512
  ---
513
 
514
+ ## 📈 Version History
515
+
516
+ ### V2 — LoKr merge onto Anima v1.0 / v1.1
517
 
518
+ - Three variants released together: **Turbo**, **Aesthetic** and **Base**
519
+ - LoKr network — linear 64 / alpha 64, conv 16 / alpha 16, full-rank, factor 4
520
+ - ~24K training steps on `anima-base-v1.0`, merged at strength 1.0
521
+ - Natural-language-only dataset; trigger words `seesee_elf` and `seesee_kitsune`
522
+ - Moves off Anima Preview3 onto the final Anima v1.0 / v1.1 releases
523
+ - Diffusion Model files — text encoder and VAE included in this repo
524
 
525
+ ### v1.0 — Initial Release
526
 
527
+ - **AniSee Base** — full fine-tune of Anima Preview3 Base
528
+ - ~20K training steps on a curated anime dataset
529
+ - LLM adapter only very lightly co-trained *(following Anima's fine-tuning guidelines)*
530
+ - Diffusion Model variant *(single `.safetensors` file)*
531
+ - Compatible with the standard Anima ComfyUI workflow
532
+ - Drop-in replacement for `anima-preview3-base.safetensors`
533
+ - Includes the official ComfyUI workflow with auto quality prefix + Qwen3-VL prompt enhancer
534
 
535
  ---
536
 
537
  ## 🔗 Links
538
 
539
+ - **CivitAI Page:** [civitai.red/models/2628747/anisee](https://civitai.red/models/2628747/anisee)
540
+ - **Example Gallery:** [anisee.anisee.workers.dev](https://anisee.anisee.workers.dev/)
541
+ - **Base Model:** [circlestone-labs/Anima](https://huggingface.co/circlestone-labs/Anima)
542
+ - **Author:** [SeeSee21 on Hugging Face](https://huggingface.co/SeeSee21)
543
 
544
  ---
545
 
546
  ## 🙏 Credits
547
 
548
+ - **Base Model:** [Anima](https://huggingface.co/circlestone-labs/Anima) by **CircleStone Labs**
549
+ and **Comfy Org** — Preview3 Base for V1, v1.0 / v1.1 for V2
550
+ - **Underlying Architecture:** Built on NVIDIA Cosmos-Predict2-2B *(Anima is a "Derivative Model")*
551
+ - **Fine-Tune and Merges:** SeeSee21
552
+ - **Workflow Custom Nodes:** yolain, ssitu, Will Miao, AILab (1038lab), rgthree
553
 
554
  ---
555
 
556
  ## 📜 License
557
 
558
+ AniSee inherits the **CircleStone Labs Non-Commercial License** from Anima. The model and
559
+ derivatives are usable **only for non-commercial purposes**. As a derivative of
560
+ Cosmos-Predict2-2B-Text2Image, the [NVIDIA Open Model License Agreement](https://www.nvidia.com/en-us/agreements/enterprise-software/nvidia-open-model-license/)
561
+ also applies insofar as it covers Derivative Models.
562
 
563
+ **Generated images are not covered by that restriction** — you may use the images you make
564
+ commercially. Selling images, paid commissions, and concept art or assets for a paid product are
565
  all fine. What needs a separate license is hosting the model behind a paid API, embedding the
566
  weights in a monetized product, or running it on a paid generation platform.
567
 
568
+ For commercial licensing of the base model, please contact CircleStone Labs at
569
+ `tdrussell@circlestone.ai`.
570
 
571
  ---
572
 
573
+ ## ❤️ Notes
574
+
575
+ AniSee is a personal anime-focused model built on Anima, made to bring a stronger anime look and
576
+ visual direction in line with my other checkpoints.
577
+
578
+ V1 was a full fine-tune of Preview3. V2 takes a different route — a LoKr merge onto the final
579
+ Anima releases, which is what makes three variants from one training run possible. Both stay
580
+ available here, because they behave differently and some people will prefer the older one.
581
+
582
+ **AniSee — clean anime, built on Anima. 🎨**
config.json CHANGED
@@ -1,177 +1,177 @@
1
- {
2
- "model_type": "anisee",
3
- "architecture": "Cosmos-Predict2-2B (Anima Derivative)",
4
- "parameters": "2B",
5
- "license": "circlestone-labs-non-commercial",
6
- "base_model": "circlestone-labs/Anima",
7
- "base_model_relation": "finetune",
8
- "author": "SeeSee21",
9
- "pipeline_tag": "text-to-image",
10
- "recommended_variant": "v2_turbo",
11
- "prompting": {
12
- "style": "tags + natural-language (mixed prompts supported)",
13
- "negative_prompt_support": "full",
14
- "recommended_prefix": "masterpiece, best quality, score_7, highres, illustration,",
15
- "aesthetic_prefix_note": "On AniSee-V2-Aesthetic, omit score_* tags in both positive and negative.",
16
- "tag_rules": {
17
- "case": "lowercase",
18
- "separator": "spaces (not underscores)",
19
- "score_tags_exception": "score_7, score_8, etc. use underscores",
20
- "artist_tags": "must be prefixed with @ (e.g. @artistname)",
21
- "weighting": "works, but needs heavier weights than SDXL (e.g. (chibi:2))"
22
- },
23
- "trigger_words": {
24
- "seesee_elf": "white-haired elf (V2 only)",
25
- "seesee_kitsune": "fox woman (V2 only)"
26
- },
27
- "trigger_note": "The V2 dataset was captioned in natural language only. Use trigger words inside descriptive sentences, not as bare tags."
28
- },
29
- "variants": {
30
- "v2_turbo": {
31
- "file": "AniSee-V2-Turbo.safetensors",
32
- "size": "~3.90 GB",
33
- "generation": "V2",
34
- "method": "LoKr merge",
35
- "merged_into": "anima-turbo-v1.1",
36
- "comfyui_path": "ComfyUI/models/diffusion_models/",
37
- "recommended_settings": {
38
- "steps": 12,
39
- "cfg": 1,
40
- "sampler": "er_sde",
41
- "scheduler": "simple",
42
- "resolution": "832x1216 (portrait, ~1 MP)"
43
- },
44
- "notes": "Fastest variant, ~6 s per image. The negative prompt has no effect at CFG 1."
45
- },
46
- "v2_aesthetic": {
47
- "file": "AniSee-V2-Aesthetic.safetensors",
48
- "size": "~3.90 GB",
49
- "generation": "V2",
50
- "method": "LoKr merge",
51
- "merged_into": "anima-aesthetic-v1.1",
52
- "comfyui_path": "ComfyUI/models/diffusion_models/",
53
- "recommended_settings": {
54
- "steps": 40,
55
- "cfg": 4.5,
56
- "sampler": "er_sde",
57
- "scheduler": "simple",
58
- "resolution": "832x1216 (portrait, ~1 MP)"
59
- },
60
- "notes": "Richest colour and softest light. Do not use score_* tags with this variant."
61
- },
62
- "v2_base": {
63
- "file": "AniSee-V2-Base.safetensors",
64
- "size": "~3.90 GB",
65
- "generation": "V2",
66
- "method": "LoKr merge",
67
- "merged_into": "anima-base-v1.0",
68
- "comfyui_path": "ComfyUI/models/diffusion_models/",
69
- "recommended_settings": {
70
- "steps": 40,
71
- "cfg": 4.5,
72
- "sampler": "er_sde",
73
- "scheduler": "simple",
74
- "resolution": "832x1216 (portrait, ~1 MP)"
75
- },
76
- "notes": "Most neutral and most flexible. The right foundation for training your own LoRAs."
77
- },
78
- "v1": {
79
- "file": "anisee.safetensors",
80
- "size": "~3.89 GB",
81
- "generation": "V1",
82
- "method": "full fine-tune",
83
- "merged_into": "anima-preview3-base",
84
- "comfyui_path": "ComfyUI/models/diffusion_models/",
85
- "recommended_settings": {
86
- "steps": 40,
87
- "cfg": 4.5,
88
- "sampler": "er_sde",
89
- "scheduler": "simple",
90
- "resolution": "832x1216 (portrait, ~1 MP)"
91
- },
92
- "notes": "Drop-in replacement for anima-preview3-base.safetensors."
93
- }
94
- },
95
- "training": {
96
- "v2": {
97
- "method": "LoKr, merged at strength 1.0",
98
- "network": {
99
- "linear": 64,
100
- "linear_alpha": 64,
101
- "conv": 16,
102
- "conv_alpha": 16,
103
- "full_rank": true,
104
- "factor": 4
105
- },
106
- "steps": "~24000",
107
- "trained_on": "anima-base-v1.0",
108
- "dataset": "curated dataset, natural-language captions only"
109
- },
110
- "v1": {
111
- "method": "full fine-tune",
112
- "steps": "~20000",
113
- "trained_on": "anima-preview3-base",
114
- "dataset": "curated anime dataset",
115
- "adapter_co_training": "lightly co-trained (per Anima guidelines)"
116
- }
117
- },
118
- "alternative_samplers": [
119
- "euler_a",
120
- "dpmpp_2m_sde_gpu",
121
- "euler"
122
- ],
123
- "planned_variants": {
124
- "v2_aio": {
125
- "description": "All-in-one V2 checkpoint (model + text encoder + VAE in one file)",
126
- "status": "planned"
127
- }
128
- },
129
- "components": {
130
- "text_encoder": {
131
- "file": "text_encoders/qwen_3_06b_base.safetensors",
132
- "description": "Anima / Qwen3 0.6B text encoder",
133
- "comfyui_path": "ComfyUI/models/text_encoders/"
134
- },
135
- "vae": {
136
- "file": "vae/qwen_image_vae.safetensors",
137
- "description": "Qwen-Image VAE (inherited from Anima)",
138
- "comfyui_path": "ComfyUI/models/vae/"
139
- }
140
- },
141
- "comfyui_paths": {
142
- "diffusion_models": "ComfyUI/models/diffusion_models/",
143
- "text_encoders": "ComfyUI/models/text_encoders/",
144
- "vae": "ComfyUI/models/vae/",
145
- "upscale_models": "ComfyUI/models/upscale_models/"
146
- },
147
- "requirements": {
148
- "custom_nodes": [
149
- "ComfyUI-Easy-Use",
150
- "ComfyUI_UltimateSDUpscale",
151
- "ComfyUI-Lora-Manager",
152
- "ComfyUI-QwenVL",
153
- "rgthree-comfy"
154
- ],
155
- "optional_upscale_model": {
156
- "file": "4x-UltraSharp.pth",
157
- "sources": [
158
- "https://openmodeldb.info/models/4x-UltraSharp",
159
- "https://huggingface.co/Kim2091/UltraSharp"
160
- ]
161
- }
162
- },
163
- "supported_vram": "8GB+",
164
- "links": {
165
- "civitai": "https://civitai.com/models/2628747/anisee",
166
- "gallery": "https://anisee.anisee.workers.dev/",
167
- "base_model": "https://huggingface.co/circlestone-labs/Anima",
168
- "author": "https://huggingface.co/SeeSee21"
169
- },
170
- "notes": [
171
- "V1 is a full fine-tune of Anima Preview3 Base. V2 is a LoKr merge onto Anima v1.0 / v1.1 — a different method, not a fine-tune.",
172
- "All checkpoints are diffusion model files: load with Load Diffusion Model, not the Checkpoint Loader.",
173
- "The Qwen text encoder and the Qwen-Image VAE are included in this repo.",
174
- "Tag rules: lowercase, spaces (not underscores), except score_7 etc. Artist tags prefixed with @.",
175
- "An all-in-one V1 checkpoint is available on CivitAI; a V2 all-in-one is planned."
176
- ]
177
- }
 
1
+ {
2
+ "model_type": "anisee",
3
+ "architecture": "Cosmos-Predict2-2B (Anima Derivative)",
4
+ "parameters": "2B",
5
+ "license": "circlestone-labs-non-commercial",
6
+ "base_model": "circlestone-labs/Anima",
7
+ "base_model_relation": "finetune",
8
+ "author": "SeeSee21",
9
+ "pipeline_tag": "text-to-image",
10
+ "recommended_variant": "v2_turbo",
11
+ "prompting": {
12
+ "style": "tags + natural-language (mixed prompts supported)",
13
+ "negative_prompt_support": "full",
14
+ "recommended_prefix": "masterpiece, best quality, score_7, highres, illustration,",
15
+ "aesthetic_prefix_note": "On AniSee-V2-Aesthetic, omit score_* tags in both positive and negative.",
16
+ "tag_rules": {
17
+ "case": "lowercase",
18
+ "separator": "spaces (not underscores)",
19
+ "score_tags_exception": "score_7, score_8, etc. use underscores",
20
+ "artist_tags": "must be prefixed with @ (e.g. @artistname)",
21
+ "weighting": "works, but needs heavier weights than SDXL (e.g. (chibi:2))"
22
+ },
23
+ "trigger_words": {
24
+ "seesee_elf": "white-haired elf (V2 only)",
25
+ "seesee_kitsune": "fox woman (V2 only)"
26
+ },
27
+ "trigger_note": "The V2 dataset was captioned in natural language only. Use trigger words inside descriptive sentences, not as bare tags."
28
+ },
29
+ "variants": {
30
+ "v2_turbo": {
31
+ "file": "AniSee-V2-Turbo.safetensors",
32
+ "size": "~3.90 GB",
33
+ "generation": "V2",
34
+ "method": "LoKr merge",
35
+ "merged_into": "anima-turbo-v1.1",
36
+ "comfyui_path": "ComfyUI/models/diffusion_models/",
37
+ "recommended_settings": {
38
+ "steps": 12,
39
+ "cfg": 1,
40
+ "sampler": "er_sde",
41
+ "scheduler": "simple",
42
+ "resolution": "832x1216 (portrait, ~1 MP)"
43
+ },
44
+ "notes": "Fastest variant, ~6 s per image. The negative prompt has no effect at CFG 1."
45
+ },
46
+ "v2_aesthetic": {
47
+ "file": "AniSee-V2-Aesthetic.safetensors",
48
+ "size": "~3.90 GB",
49
+ "generation": "V2",
50
+ "method": "LoKr merge",
51
+ "merged_into": "anima-aesthetic-v1.1",
52
+ "comfyui_path": "ComfyUI/models/diffusion_models/",
53
+ "recommended_settings": {
54
+ "steps": 40,
55
+ "cfg": 4.5,
56
+ "sampler": "er_sde",
57
+ "scheduler": "simple",
58
+ "resolution": "832x1216 (portrait, ~1 MP)"
59
+ },
60
+ "notes": "Richest colour and softest light. Do not use score_* tags with this variant."
61
+ },
62
+ "v2_base": {
63
+ "file": "AniSee-V2-Base.safetensors",
64
+ "size": "~3.90 GB",
65
+ "generation": "V2",
66
+ "method": "LoKr merge",
67
+ "merged_into": "anima-base-v1.0",
68
+ "comfyui_path": "ComfyUI/models/diffusion_models/",
69
+ "recommended_settings": {
70
+ "steps": 40,
71
+ "cfg": 4.5,
72
+ "sampler": "er_sde",
73
+ "scheduler": "simple",
74
+ "resolution": "832x1216 (portrait, ~1 MP)"
75
+ },
76
+ "notes": "Most neutral and most flexible. The right foundation for training your own LoRAs."
77
+ },
78
+ "v1": {
79
+ "file": "anisee.safetensors",
80
+ "size": "~3.89 GB",
81
+ "generation": "V1",
82
+ "method": "full fine-tune",
83
+ "merged_into": "anima-preview3-base",
84
+ "comfyui_path": "ComfyUI/models/diffusion_models/",
85
+ "recommended_settings": {
86
+ "steps": 40,
87
+ "cfg": 4.5,
88
+ "sampler": "er_sde",
89
+ "scheduler": "simple",
90
+ "resolution": "832x1216 (portrait, ~1 MP)"
91
+ },
92
+ "notes": "Drop-in replacement for anima-preview3-base.safetensors."
93
+ }
94
+ },
95
+ "training": {
96
+ "v2": {
97
+ "method": "LoKr, merged at strength 1.0",
98
+ "network": {
99
+ "linear": 64,
100
+ "linear_alpha": 64,
101
+ "conv": 16,
102
+ "conv_alpha": 16,
103
+ "full_rank": true,
104
+ "factor": 4
105
+ },
106
+ "steps": "~24000",
107
+ "trained_on": "anima-base-v1.0",
108
+ "dataset": "curated dataset, natural-language captions only"
109
+ },
110
+ "v1": {
111
+ "method": "full fine-tune",
112
+ "steps": "~20000",
113
+ "trained_on": "anima-preview3-base",
114
+ "dataset": "curated anime dataset",
115
+ "adapter_co_training": "lightly co-trained (per Anima guidelines)"
116
+ }
117
+ },
118
+ "alternative_samplers": [
119
+ "euler_a",
120
+ "dpmpp_2m_sde_gpu",
121
+ "euler"
122
+ ],
123
+ "planned_variants": {
124
+ "v2_aio": {
125
+ "description": "All-in-one V2 checkpoint (model + text encoder + VAE in one file)",
126
+ "status": "planned"
127
+ }
128
+ },
129
+ "components": {
130
+ "text_encoder": {
131
+ "file": "text_encoders/qwen_3_06b_base.safetensors",
132
+ "description": "Anima / Qwen3 0.6B text encoder",
133
+ "comfyui_path": "ComfyUI/models/text_encoders/"
134
+ },
135
+ "vae": {
136
+ "file": "vae/qwen_image_vae.safetensors",
137
+ "description": "Qwen-Image VAE (inherited from Anima)",
138
+ "comfyui_path": "ComfyUI/models/vae/"
139
+ }
140
+ },
141
+ "comfyui_paths": {
142
+ "diffusion_models": "ComfyUI/models/diffusion_models/",
143
+ "text_encoders": "ComfyUI/models/text_encoders/",
144
+ "vae": "ComfyUI/models/vae/",
145
+ "upscale_models": "ComfyUI/models/upscale_models/"
146
+ },
147
+ "requirements": {
148
+ "custom_nodes": [
149
+ "ComfyUI-Easy-Use",
150
+ "ComfyUI_UltimateSDUpscale",
151
+ "ComfyUI-Lora-Manager",
152
+ "ComfyUI-QwenVL",
153
+ "rgthree-comfy"
154
+ ],
155
+ "optional_upscale_model": {
156
+ "file": "4x-UltraSharp.pth",
157
+ "sources": [
158
+ "https://openmodeldb.info/models/4x-UltraSharp",
159
+ "https://huggingface.co/Kim2091/UltraSharp"
160
+ ]
161
+ }
162
+ },
163
+ "supported_vram": "8GB+",
164
+ "links": {
165
+ "civitai": "https://civitai.red/models/2628747/anisee",
166
+ "gallery": "https://anisee.anisee.workers.dev/",
167
+ "base_model": "https://huggingface.co/circlestone-labs/Anima",
168
+ "author": "https://huggingface.co/SeeSee21"
169
+ },
170
+ "notes": [
171
+ "V1 is a full fine-tune of Anima Preview3 Base. V2 is a LoKr merge onto Anima v1.0 / v1.1 — a different method, not a fine-tune.",
172
+ "All checkpoints are diffusion model files: load with Load Diffusion Model, not the Checkpoint Loader.",
173
+ "The Qwen text encoder and the Qwen-Image VAE are included in this repo.",
174
+ "Tag rules: lowercase, spaces (not underscores), except score_7 etc. Artist tags prefixed with @.",
175
+ "An all-in-one V1 checkpoint is available on CivitAI; a V2 all-in-one is planned."
176
+ ]
177
+ }