2.58 GB
20 files
Updated about 1 month ago
Name
Size
data
.gitattributes2.5 kB
xet
README.md10 kB
xet
manifest.json5.87 kB
xet
metadata.csv2.75 MB
xet
README.md

Dataset.ET Amharic Speech — v0.1.0

22.706 hours · 7,405 clips · 320 speakers · 7,145 distinct prompts

Dataset Description

Dataset Summary

Read speech in Amharic, crowdsourced from volunteer contributors in Ethiopia through a Telegram bot, peer-validated by other contributors, and screened acoustically before release. Amharic has very little open speech data; this corpus exists to change that.

Contributors read a displayed prompt aloud, other contributors listen and vote on whether the recording matches, and contributors earn mobile airtime for accepted work. The corpus is built by the community it serves, and published by Snapwre Technologies PLC.

Supported Tasks

  • automatic-speech-recognition — the primary use. Each clip pairs audio with the exact prompt text that was read.
  • Speaker and demographic analysis, within the limits described below.

Languages

Amharic (am), written in the Ge'ez script.

Dataset Structure

Data Instances

from datasets import load_dataset
ds = load_dataset("snapwre/amharic-speech", split="train")
ds[0]
{
  "audio": {"array": array([...]), "sampling_rate": 16000},
  "sentence": "…",
  "speaker_id": "spk_1a2b3c4d5e6f",
  "duration_s": 9.29,
  "speech_start_s": 0.31, "speech_end_s": 9.29,
  "gender": "male", "age_band": "18-24", "region": "addis_ababa",
  ...
}

Data Fields

Field Type Description
audio Audio(16 kHz) 16 kHz mono FLAC
clip_id string Release-local identifier
sentence string The prompt the contributor was asked to read
speaker_id string Pseudonymous, salted per release; not linkable across releases
language string ISO code
duration_s float32 Total clip duration, measured from the audio
speech_s float32 Duration excluding detected silence
speech_start_s, speech_end_s float32 Where speech begins and ends. Audio is not trimmed — slice it yourself if you want silence removed
lufs float32 Integrated loudness (EBU R128)
gender, age_band, region string Self-reported, optional. null may mean not stated or withheld for privacy — see Personal and Sensitive Information
sample_rate int32 Always 16000
up_votes, down_votes int16 Peer validation votes

Data Splits

Splits are speaker-disjoint and prompt-disjoint. No voice and no sentence appears in more than one split, so evaluation is not inflated by a model recognising a speaker or having memorised a prompt.

Split Clips Hours Speakers Prompts
test 718 2.222 34 715
validation 707 2.157 28 704
train 5,980 18.326 258 5,733

metadata.csv carries every field except the audio, for inspecting the corpus without downloading it.

Dataset Creation

Curation Rationale

Selected conservatively: a first release that holds up matters more than a large one. Only clips with unanimous peer approval were considered, and those were then screened acoustically.

Source Data

Recordings are original contributions captured through the Dataset.ET Telegram bot as Opus voice messages, losslessly transcoded to 16 kHz mono FLAC. Each was made by a contributor reading a displayed prompt aloud.

Annotations

The prompt text is the transcript — contributors read a known sentence rather than transcribing free speech, so transcripts are exact by construction. Peer validators voted on whether each recording matched its prompt.

Personal and Sensitive Information

Speech is biometric. Contributors granted an explicit licence to publish their recordings in open datasets, and that consent is the basis for this release. Everything published alongside the voice has been treated as follows:

  • Database identifiers never ship. speaker_id and clip_id are salted hashes, and the salt is not published.
  • Demographic fields satisfy k-anonymity at k=5: every published combination of gender, age band and region describes at least five contributors. Rarer combinations were suppressed to null, so a null may mean withheld rather than not stated.
  • Prompt text was scanned for phone numbers, email addresses and URLs.

Residual risk worth stating plainly: no automated check has verified that no contributor spoke identifying information aloud within a recording.

Considerations for Using the Data

Discussion of Biases

Gender Clips Share
male 5,062 68.4%
female 2,303 31.1%
(not stated) 40 0.5%
Age band Clips Share
18-24 5,559 75.1%
25-34 1,457 19.7%
(not stated) 251 3.4%
35-44 138 1.9%
Region Clips Share
addis_ababa 3,514 47.5%
amhara 1,805 24.4%
(not stated) 861 11.6%
oromia 566 7.6%
other 363 4.9%
central_ethiopia 183 2.5%
sidama 113 1.5%

Contributors are overwhelmingly young and urban. Models trained on this will perform worse on older and rural speakers. Do not treat this as a representative sample of Amharic speakers.

Other Known Limitations

  • Read speech, not conversation. Prosody, disfluency and turn-taking are unlike spontaneous speech.
  • Narrow prompt domain. Prompt vocabulary skews formal and toward current affairs, so coverage of casual and conversational registers is thin.
  • Telegram Opus origin. Captured through Telegram voice messages, carrying that codec's artefacts and whatever processing contributors' devices applied. Not studio audio.
  • Not loudness-normalised. Normalisation is a training-time choice; the lufs column lets you do it deterministically.
  • Peer validation is imperfect. During the period this data was collected, validators were paid per validation rather than per correct validation, which rewards approving quickly. Clips with any reject vote were excluded and the acoustic screening below exists to compensate, but expect residual noise.

Selection and Screening

Starting from every clip that passed peer validation, these were excluded:

Reason Clips
not needed to reach target hours 23,429
fewer than 3 accept votes 16,629
shorter than 1.5s 5,435
duplicate audio 661
contested (at least one reject vote) 467
speaker below minimum clip count 373
longer than 30.0s 210

Survivors were screened acoustically:

Rejected by screening Clips
prompt not fully read 840
too quiet to recover 334
clipped 333
mostly silence 331
almost no speech 286

Additional Information

Licensing Information

Audio: CC BY 4.0. Contributors granted Dataset.ET a perpetual, irrevocable, worldwide, royalty-free licence to publish and distribute their recordings as part of open datasets, and consented to release under a permissive open data licence. Those rights are held by Snapwre Technologies PLC, which operates Dataset.ET, and are exercised here to publish the corpus openly.

Prompt text: the sentences contributors read are third-party material and are reproduced here as transcripts of the recordings. Snapwre Technologies PLC makes no licence grant over the prompt text itself; anyone intending to redistribute the text separately from the audio should assess that independently.

Citation Information

@misc{datasetet_am_0_1_0,
  title  = {Dataset.ET Amharic Speech v0.1.0},
  author = {Dataset.ET contributors and Snapwre Technologies PLC},
  publisher = {Snapwre Technologies PLC},
  year   = {2026},
  url    = {https://huggingface.co/datasets/snapwre/amharic-speech}
}

Dataset Curators

Dataset.ET is a project of Snapwre Technologies PLC, Addis Ababa, Ethiopia. The company operates the collection platform, holds the contributor-granted rights to the recordings, and publishes the corpus openly.

Contributions

Built by 320 volunteer contributors in Ethiopia, plus the validators who reviewed their recordings. The corpus exists because they gave their voices to it.

Every shard's SHA-256, and the exact selection, screening and anonymisation policies that produced this release, are recorded in manifest.json.

Total size
2.58 GB
Files
20
Last updated
Aug 26
Pre-warmed CDN
US EU US EU

Contributors