Buckets:
| Name | Size | Uploaded | Xet hash |
|---|---|---|---|
| data | 16 items | ||
| .gitattributes | 2.5 kB xet | 738f1125 | |
| README.md | 10 kB xet | a2db8c64 | |
| manifest.json | 5.87 kB xet | d2d6478e | |
| metadata.csv | 2.75 MB xet | eee0fe81 |
Dataset.ET Amharic Speech — v0.1.0
22.706 hours · 7,405 clips · 320 speakers · 7,145 distinct prompts
Dataset Description
- Curated by: Snapwre Technologies PLC (Addis Ababa, Ethiopia)
- Homepage: https://dataset.et
- Repository: https://github.com/snapwre/dataset-et-release
- Point of Contact: https://huggingface.co/snapwre
Dataset Summary
Read speech in Amharic, crowdsourced from volunteer contributors in Ethiopia through a Telegram bot, peer-validated by other contributors, and screened acoustically before release. Amharic has very little open speech data; this corpus exists to change that.
Contributors read a displayed prompt aloud, other contributors listen and vote on whether the recording matches, and contributors earn mobile airtime for accepted work. The corpus is built by the community it serves, and published by Snapwre Technologies PLC.
Supported Tasks
automatic-speech-recognition— the primary use. Each clip pairs audio with the exact prompt text that was read.- Speaker and demographic analysis, within the limits described below.
Languages
Amharic (am), written in the Ge'ez script.
Dataset Structure
Data Instances
from datasets import load_dataset
ds = load_dataset("snapwre/amharic-speech", split="train")
ds[0]
{
"audio": {"array": array([...]), "sampling_rate": 16000},
"sentence": "…",
"speaker_id": "spk_1a2b3c4d5e6f",
"duration_s": 9.29,
"speech_start_s": 0.31, "speech_end_s": 9.29,
"gender": "male", "age_band": "18-24", "region": "addis_ababa",
...
}
Data Fields
| Field | Type | Description |
|---|---|---|
audio |
Audio(16 kHz) | 16 kHz mono FLAC |
clip_id |
string | Release-local identifier |
sentence |
string | The prompt the contributor was asked to read |
speaker_id |
string | Pseudonymous, salted per release; not linkable across releases |
language |
string | ISO code |
duration_s |
float32 | Total clip duration, measured from the audio |
speech_s |
float32 | Duration excluding detected silence |
speech_start_s, speech_end_s |
float32 | Where speech begins and ends. Audio is not trimmed — slice it yourself if you want silence removed |
lufs |
float32 | Integrated loudness (EBU R128) |
gender, age_band, region |
string | Self-reported, optional. null may mean not stated or withheld for privacy — see Personal and Sensitive Information |
sample_rate |
int32 | Always 16000 |
up_votes, down_votes |
int16 | Peer validation votes |
Data Splits
Splits are speaker-disjoint and prompt-disjoint. No voice and no sentence appears in more than one split, so evaluation is not inflated by a model recognising a speaker or having memorised a prompt.
| Split | Clips | Hours | Speakers | Prompts |
|---|---|---|---|---|
| test | 718 | 2.222 | 34 | 715 |
| validation | 707 | 2.157 | 28 | 704 |
| train | 5,980 | 18.326 | 258 | 5,733 |
metadata.csv carries every field except the audio, for inspecting the corpus
without downloading it.
Dataset Creation
Curation Rationale
Selected conservatively: a first release that holds up matters more than a large one. Only clips with unanimous peer approval were considered, and those were then screened acoustically.
Source Data
Recordings are original contributions captured through the Dataset.ET Telegram bot as Opus voice messages, losslessly transcoded to 16 kHz mono FLAC. Each was made by a contributor reading a displayed prompt aloud.
Annotations
The prompt text is the transcript — contributors read a known sentence rather than transcribing free speech, so transcripts are exact by construction. Peer validators voted on whether each recording matched its prompt.
Personal and Sensitive Information
Speech is biometric. Contributors granted an explicit licence to publish their recordings in open datasets, and that consent is the basis for this release. Everything published alongside the voice has been treated as follows:
- Database identifiers never ship.
speaker_idandclip_idare salted hashes, and the salt is not published. - Demographic fields satisfy k-anonymity at k=5: every published combination
of gender, age band and region describes at least five contributors. Rarer
combinations were suppressed to
null, so anullmay mean withheld rather than not stated. - Prompt text was scanned for phone numbers, email addresses and URLs.
Residual risk worth stating plainly: no automated check has verified that no contributor spoke identifying information aloud within a recording.
Considerations for Using the Data
Discussion of Biases
| Gender | Clips | Share |
|---|---|---|
| male | 5,062 | 68.4% |
| female | 2,303 | 31.1% |
| (not stated) | 40 | 0.5% |
| Age band | Clips | Share |
|---|---|---|
| 18-24 | 5,559 | 75.1% |
| 25-34 | 1,457 | 19.7% |
| (not stated) | 251 | 3.4% |
| 35-44 | 138 | 1.9% |
| Region | Clips | Share |
|---|---|---|
| addis_ababa | 3,514 | 47.5% |
| amhara | 1,805 | 24.4% |
| (not stated) | 861 | 11.6% |
| oromia | 566 | 7.6% |
| other | 363 | 4.9% |
| central_ethiopia | 183 | 2.5% |
| sidama | 113 | 1.5% |
Contributors are overwhelmingly young and urban. Models trained on this will perform worse on older and rural speakers. Do not treat this as a representative sample of Amharic speakers.
Other Known Limitations
- Read speech, not conversation. Prosody, disfluency and turn-taking are unlike spontaneous speech.
- Narrow prompt domain. Prompt vocabulary skews formal and toward current affairs, so coverage of casual and conversational registers is thin.
- Telegram Opus origin. Captured through Telegram voice messages, carrying that codec's artefacts and whatever processing contributors' devices applied. Not studio audio.
- Not loudness-normalised. Normalisation is a training-time choice; the
lufscolumn lets you do it deterministically. - Peer validation is imperfect. During the period this data was collected, validators were paid per validation rather than per correct validation, which rewards approving quickly. Clips with any reject vote were excluded and the acoustic screening below exists to compensate, but expect residual noise.
Selection and Screening
Starting from every clip that passed peer validation, these were excluded:
| Reason | Clips |
|---|---|
| not needed to reach target hours | 23,429 |
| fewer than 3 accept votes | 16,629 |
| shorter than 1.5s | 5,435 |
| duplicate audio | 661 |
| contested (at least one reject vote) | 467 |
| speaker below minimum clip count | 373 |
| longer than 30.0s | 210 |
Survivors were screened acoustically:
| Rejected by screening | Clips |
|---|---|
| prompt not fully read | 840 |
| too quiet to recover | 334 |
| clipped | 333 |
| mostly silence | 331 |
| almost no speech | 286 |
Additional Information
Licensing Information
Audio: CC BY 4.0. Contributors granted Dataset.ET a perpetual, irrevocable, worldwide, royalty-free licence to publish and distribute their recordings as part of open datasets, and consented to release under a permissive open data licence. Those rights are held by Snapwre Technologies PLC, which operates Dataset.ET, and are exercised here to publish the corpus openly.
Prompt text: the sentences contributors read are third-party material and are reproduced here as transcripts of the recordings. Snapwre Technologies PLC makes no licence grant over the prompt text itself; anyone intending to redistribute the text separately from the audio should assess that independently.
Citation Information
@misc{datasetet_am_0_1_0,
title = {Dataset.ET Amharic Speech v0.1.0},
author = {Dataset.ET contributors and Snapwre Technologies PLC},
publisher = {Snapwre Technologies PLC},
year = {2026},
url = {https://huggingface.co/datasets/snapwre/amharic-speech}
}
Dataset Curators
Dataset.ET is a project of Snapwre Technologies PLC, Addis Ababa, Ethiopia. The company operates the collection platform, holds the contributor-granted rights to the recordings, and publishes the corpus openly.
Contributions
Built by 320 volunteer contributors in Ethiopia, plus the validators who reviewed their recordings. The corpus exists because they gave their voices to it.
Every shard's SHA-256, and the exact selection, screening and anonymisation
policies that produced this release, are recorded in manifest.json.
- Total size
- 2.58 GB
- Files
- 20
- Last updated
- Aug 26
- Pre-warmed CDN
- US EU US EU