Genius model! - and...improvement idea!
Hey bench-labs Team!
I saw this model and I am VERY impressed by your work!! π₯
As I read, you've used the ~3M images recaptioned dataset.
Here's another dataset that recaptioned the dataset:
https://huggingface.co/datasets/opendiffusionai/cc12m-cleaned
It's 8.5M rows! I hope it's useful for you.
Best regards,
LH-Tech-AI
Hey LH-Tech-AI, thanks so much for the kind words and for flagging this really appreciate it!
cc12m-cleaned looks like a solid resource, especially with 8.5M rows. We'll take a look and see how it fits with our current data pipeline / recaptioning setup. Appreciate you taking the time to share it!
Hey, no problem
Have fun
Is a v6 expected to come soon? π₯
Is a v6 expected to come soon? π₯
yep, its in the pretraining stage right now with a new architecture using a joint MMDiT-style attention transformer (image and text tokens attend to each other directly, not just cross-attention), conditioned on both T5-base and CLIP text embeddings, with 2D rotary position embeddings, QK-norm, and SwiGLU MLPs, plus a REPA loss that aligns the model's mid-layer features with frozen DINOv2 to help it learn structure faster. It runs on an upgraded SDXL VAE, at about 155M trainable parameters, roughly four times v5's size.
Nice! Can you give early samples please? π€
Okay, thank you π€
This looks great for early stage of training.
On how many images and how many epochs are you training and when do you expect it to be released?
currently 2.8M with a planned v6.1 using https://huggingface.co/datasets/opendiffusionai/cc12m-cleaned
Alright. Sounds good.
Can you give me new samples please? I am hyped to see this model!
Ok, thanks!
nice! thanks

