Datasets, selected speech ladder, voice-acting training views, predictors and generation model. Hours overlap; see source cards.
LAION eV
non-profit
AI & ML interests
open multi-modal foundation models and datasets for their creation; scaling laws, model evaluation; fully local, sovereign model deployment, personalized assistants and open local agentic systems
Recent Activity
View all activity
Timed-script, preference-pair and synthetic ranked-take releases. Audio overlaps; not all views trained SFT3.
Annotated English/German speech. Repository-view hours overlap; audio and annotation rights vary by source. See each dataset card.
A 10-Million-Hour Open Video Dataset for Multimodal Pre-training
Collection of models and dataset related to MixtureVitae, open and fully reproducible pretraining dataset built from permissive sources
The full collection of our EmoNet effort. More info available at: https://huggingface.co/blog/felfri/emonet
Releases related to Open-ψ (Open-Sci) Collective
Re-LAION-5B-research
OpenCLIP models trained on DataComp (https://huggingface.co/papers/2304.14108).
-
laion/CLIP-ViT-L-14-DataComp.XL-s13B-b90K
Zero-Shot Image Classification • Updated • 40.7k • 127 -
laion/CLIP-ViT-B-16-DataComp.XL-s13B-b90K
Zero-Shot Image Classification • Updated • 13.5k • 10 -
laion/CLIP-ViT-B-32-256x256-DataComp-s34B-b86K
Zero-Shot Image Classification • Updated • 2.75k • 9 -
laion/CLIP-ViT-B-16-DataComp.L-s1B-b8K
Zero-Shot Image Classification • Updated • 338 • 2
CLAP is to audio what CLIP is to image.
-
laion/larger_clap_general
Feature Extraction • Updated • 269k • • 54 -
laion/larger_clap_music_and_speech
Feature Extraction • Updated • 316k • • 46 -
laion/larger_clap_music
Feature Extraction • Updated • 49.4k • • 50 -
laion/clap-htsat-fused
Audio Classification • 0.2B • Updated • 7.09M • 150
Genuineness and vocal-burst-blend predictors, plus SFT3. Both predictor repositories bundle the frozen VoiceCLAP encoder.
Nested DE/EN training views: 20,001.5, 49,569.2 and 98,920.1 h; separate 80.1 h holdout. Source-specific audio rights apply.
RL datasets (mostly agentic) that were validated with oracle solutions. Validation protocol: https://gist.github.com/marianna13/f9bc94d2ecb39c302c6c9b
27 Delphi midtraining endpoint checkpoints (K=0.20): 9 scales x 3 mixes. Dense Qwen3, Llama-3 tok. marin#6279.
models and datasets related to openthoughts 4 experiments
-
laion/openthoughts-4-code-qwen3-32b-annotated-32k_qwen3-1.7B_32k
2B • Updated • 23 -
laion/openthoughts-4-code-qwen3-32b-annotated-32k_qwen2.5-1.5B_32k
Text Generation • 2B • Updated • 72 • 1 -
laion/openthoughts-3-QwQ-32b-annotated-16k_qwen2.5-1.5B_16k
Text Generation • 2B • Updated • 37 -
laion/openthoughts-4-code-qwen3-32b-annotated-7k_qwen3-1.7B_10k
Text Generation • 2B • Updated • 18
openMaMMUT/openCLIP models trained on DataComp-1.4B, DFN-1.4B and Re-LAION-2B. Pre-trained models on various scales, incl. intermediate checkpoints
-
laion/openMaMMUT-ViT-L-14-DataComp-1.4B-s12.8B-b180K
Zero-Shot Image Classification • Updated • 17 • 6 -
Scaling Laws for Robust Comparison of Open Foundation Language-Vision Models and Datasets
Paper • 2506.04598 • Published • 7 -
laion/openMaMMUT-ViT-L-14-512x512-pt_datacomp1b-ft_DFN512x512-s293M-b32k
Zero-Shot Image Classification • Updated • 16 • 2 -
laion/scaling-laws-for-comparison
Updated • 3
Re-LAION-5B research safe
OpenCLIP models trained on LAION-2B
-
laion/CLIP-ViT-bigG-14-laion2B-39B-b160k
Zero-Shot Image Classification • Updated • 88.5k • 321 -
laion/CLIP-ViT-g-14-laion2B-s34B-b88K
Zero-Shot Image Classification • Updated • 4.71k • 29 -
laion/CLIP-ViT-g-14-laion2B-s12B-b42K
1B • Updated • 17.4k • 44 -
laion/CLIP-ViT-H-14-laion2B-s32B-b79K
Zero-Shot Image Classification • 1.0B • Updated • 616k • 483
Datasets, selected speech ladder, voice-acting training views, predictors and generation model. Hours overlap; see source cards.
Genuineness and vocal-burst-blend predictors, plus SFT3. Both predictor repositories bundle the frozen VoiceCLAP encoder.
Timed-script, preference-pair and synthetic ranked-take releases. Audio overlaps; not all views trained SFT3.
Nested DE/EN training views: 20,001.5, 49,569.2 and 98,920.1 h; separate 80.1 h holdout. Source-specific audio rights apply.
Annotated English/German speech. Repository-view hours overlap; audio and annotation rights vary by source. See each dataset card.
RL datasets (mostly agentic) that were validated with oracle solutions. Validation protocol: https://gist.github.com/marianna13/f9bc94d2ecb39c302c6c9b
A 10-Million-Hour Open Video Dataset for Multimodal Pre-training
27 Delphi midtraining endpoint checkpoints (K=0.20): 9 scales x 3 mixes. Dense Qwen3, Llama-3 tok. marin#6279.
Collection of models and dataset related to MixtureVitae, open and fully reproducible pretraining dataset built from permissive sources
models and datasets related to openthoughts 4 experiments
-
laion/openthoughts-4-code-qwen3-32b-annotated-32k_qwen3-1.7B_32k
2B • Updated • 23 -
laion/openthoughts-4-code-qwen3-32b-annotated-32k_qwen2.5-1.5B_32k
Text Generation • 2B • Updated • 72 • 1 -
laion/openthoughts-3-QwQ-32b-annotated-16k_qwen2.5-1.5B_16k
Text Generation • 2B • Updated • 37 -
laion/openthoughts-4-code-qwen3-32b-annotated-7k_qwen3-1.7B_10k
Text Generation • 2B • Updated • 18
The full collection of our EmoNet effort. More info available at: https://huggingface.co/blog/felfri/emonet
Releases related to Open-ψ (Open-Sci) Collective
openMaMMUT/openCLIP models trained on DataComp-1.4B, DFN-1.4B and Re-LAION-2B. Pre-trained models on various scales, incl. intermediate checkpoints
-
laion/openMaMMUT-ViT-L-14-DataComp-1.4B-s12.8B-b180K
Zero-Shot Image Classification • Updated • 17 • 6 -
Scaling Laws for Robust Comparison of Open Foundation Language-Vision Models and Datasets
Paper • 2506.04598 • Published • 7 -
laion/openMaMMUT-ViT-L-14-512x512-pt_datacomp1b-ft_DFN512x512-s293M-b32k
Zero-Shot Image Classification • Updated • 16 • 2 -
laion/scaling-laws-for-comparison
Updated • 3
Re-LAION-5B-research
Re-LAION-5B research safe
OpenCLIP models trained on DataComp (https://huggingface.co/papers/2304.14108).
-
laion/CLIP-ViT-L-14-DataComp.XL-s13B-b90K
Zero-Shot Image Classification • Updated • 40.7k • 127 -
laion/CLIP-ViT-B-16-DataComp.XL-s13B-b90K
Zero-Shot Image Classification • Updated • 13.5k • 10 -
laion/CLIP-ViT-B-32-256x256-DataComp-s34B-b86K
Zero-Shot Image Classification • Updated • 2.75k • 9 -
laion/CLIP-ViT-B-16-DataComp.L-s1B-b8K
Zero-Shot Image Classification • Updated • 338 • 2
OpenCLIP models trained on LAION-2B
-
laion/CLIP-ViT-bigG-14-laion2B-39B-b160k
Zero-Shot Image Classification • Updated • 88.5k • 321 -
laion/CLIP-ViT-g-14-laion2B-s34B-b88K
Zero-Shot Image Classification • Updated • 4.71k • 29 -
laion/CLIP-ViT-g-14-laion2B-s12B-b42K
1B • Updated • 17.4k • 44 -
laion/CLIP-ViT-H-14-laion2B-s32B-b79K
Zero-Shot Image Classification • 1.0B • Updated • 616k • 483
CLAP is to audio what CLIP is to image.
-
laion/larger_clap_general
Feature Extraction • Updated • 269k • • 54 -
laion/larger_clap_music_and_speech
Feature Extraction • Updated • 316k • • 46 -
laion/larger_clap_music
Feature Extraction • Updated • 49.4k • • 50 -
laion/clap-htsat-fused
Audio Classification • 0.2B • Updated • 7.09M • 150