Current release standard: spontaneous speech, human-validated transcripts, word-level forced alignment. Built for evaluation, not training.
AI & ML interests
Speech data for underrepresented languages, accents and niche domains, at scale. 2.5M+ consented contributors | 180+ countries | ~350 languages collectable | ~500,000 hours off the shelf 📊 Proprietary, first-party recordings with consent and provenance records, not available in any other dataset. Human-validated transcription for every language, tailored to your needs. Anything not off the shelf, sourced through our community. 📧 https://www.silencio.network/contact
Recent Activity
English, French and Spanish as spoken worldwide: speakers from dozens of countries of birth, labelled by first language, accent and proficiency.
Philippine-language speech under one protocol. 7,500 hours in active collection across Hiligaynon, Tagalog and Cebuano.
Spontaneous African-language speech with human-validated transcription and word-level alignment. Samples; the catalogue behind them is far larger.
Spontaneous speech from South Asia: Hindi, Urdu, Bengali and more, one clip per speaker, with regional variety and speaker metadata.
Accented English, global French and clinical-domain speech. Transcript coverage stated per card; human transcription available on request.
Current release standard: spontaneous speech, human-validated transcripts, word-level forced alignment. Built for evaluation, not training.
Spontaneous African-language speech with human-validated transcription and word-level alignment. Samples; the catalogue behind them is far larger.
English, French and Spanish as spoken worldwide: speakers from dozens of countries of birth, labelled by first language, accent and proficiency.
Spontaneous speech from South Asia: Hindi, Urdu, Bengali and more, one clip per speaker, with regional variety and speaker metadata.
Philippine-language speech under one protocol. 7,500 hours in active collection across Hiligaynon, Tagalog and Cebuano.
Accented English, global French and clinical-domain speech. Transcript coverage stated per card; human transcription available on request.