AI & ML interests

open multi-modal foundation models and datasets for their creation; scaling laws, model evaluation; fully local, sovereign model deployment, personalized assistants and open local agentic systems

Recent Activity

ChristophSchuhmann  updated a collection about 1 hour ago
Humaneness Voice: Research Artifacts
ChristophSchuhmann  updated a collection about 1 hour ago
Humaneness Voice: Research Artifacts
ChristophSchuhmann  updated a collection about 1 hour ago
Humaneness Voice: Research Artifacts
View all activity

laion 's collections 19

Humaneness Voice: Research Artifacts
Datasets, selected speech ladder, voice-acting training views, predictors and generation model. Hours overlap; see source cards.
Voice Acting: Training and Preference Views
Timed-script, preference-pair and synthetic ranked-take releases. Audio overlaps; not all views trained SFT3.
Expressive Speech: Annotated Sources
Annotated English/German speech. Repository-view hours overlap; audio and annotation rights vary by source. See each dataset card.
MixtureVitae study models and datasets
Collection of models and dataset related to MixtureVitae, open and fully reproducible pretraining dataset built from permissive sources
EmoNet
The full collection of our EmoNet effort. More info available at: https://huggingface.co/blog/felfri/emonet
Open-ψ (Open-Sci)
Releases related to Open-ψ (Open-Sci) Collective
OpenCLIP DataComp
OpenCLIP models trained on DataComp (https://huggingface.co/papers/2304.14108).
CLAP: Contrastive Language-Audio Pretraining
CLAP is to audio what CLIP is to image.
OT Validated RL datasets
RL datasets (mostly agentic) that were validated with oracle solutions. Validation protocol: https://gist.github.com/marianna13/f9bc94d2ecb39c302c6c9b
openthoughts 4 experiments
models and datasets related to openthoughts 4 experiments
openMaMMUT/openCLIP models and scaling laws
openMaMMUT/openCLIP models trained on DataComp-1.4B, DFN-1.4B and Re-LAION-2B. Pre-trained models on various scales, incl. intermediate checkpoints
OpenCLIP LAION-2B
OpenCLIP models trained on LAION-2B
Humaneness Voice: Research Artifacts
Datasets, selected speech ladder, voice-acting training views, predictors and generation model. Hours overlap; see source cards.
Voice Acting: Training and Preference Views
Timed-script, preference-pair and synthetic ranked-take releases. Audio overlaps; not all views trained SFT3.
Expressive Speech: Annotated Sources
Annotated English/German speech. Repository-view hours overlap; audio and annotation rights vary by source. See each dataset card.
OT Validated RL datasets
RL datasets (mostly agentic) that were validated with oracle solutions. Validation protocol: https://gist.github.com/marianna13/f9bc94d2ecb39c302c6c9b
MixtureVitae study models and datasets
Collection of models and dataset related to MixtureVitae, open and fully reproducible pretraining dataset built from permissive sources
openthoughts 4 experiments
models and datasets related to openthoughts 4 experiments
EmoNet
The full collection of our EmoNet effort. More info available at: https://huggingface.co/blog/felfri/emonet
Open-ψ (Open-Sci)
Releases related to Open-ψ (Open-Sci) Collective
openMaMMUT/openCLIP models and scaling laws
openMaMMUT/openCLIP models trained on DataComp-1.4B, DFN-1.4B and Re-LAION-2B. Pre-trained models on various scales, incl. intermediate checkpoints
OpenCLIP DataComp
OpenCLIP models trained on DataComp (https://huggingface.co/papers/2304.14108).
OpenCLIP LAION-2B
OpenCLIP models trained on LAION-2B
CLAP: Contrastive Language-Audio Pretraining
CLAP is to audio what CLIP is to image.