Instructions to use drivendata/pasketti-phonetic with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use drivendata/pasketti-phonetic with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("automatic-speech-recognition", model="drivendata/pasketti-phonetic")# pip install -U transformers accelerate # Load model directly from transformers import AutoProcessor, AutoModelForCTC processor = AutoProcessor.from_pretrained("drivendata/pasketti-phonetic") model = AutoModelForCTC.from_pretrained("drivendata/pasketti-phonetic", device_map="auto") - Notebooks
- Google Colab
- Kaggle
pasketti-phonetic 🍝
pasketti-phonetic is an automatic speech recognition (ASR) model designed specifically for child speech. pasketti-phonetic generates phonetic (per-sound) transcriptions in the International Phonetic Alphabet (IPA), including nasalization and diphthong markings, from audio files.
🐣 Existing open-source ASR models trained on adult speech perform poorly on children because child speech is fundamentally different from adult speech. Children are still learning to produce speech sounds and developing fine motor skills. Common child speech errors include:
- Metathesis: "elephant" → "ephelant"
- Velar fronting: "cup" → "tup"
- Syllable deletion: "banana" → "nana"
- Combinations of errors: "spaghetti" → "pasketti" 🍝
🔉 pasketti-phonetic is fine-tuned on a large corpus of 116 hours of phonetically transcribed child speech spanning roughly 1,700 different children. This is the largest dataset of its kind for use in developing and testing open ASR models for kids. The model is designed to advance early education assessments, teaching tools, and needs screening.
📈 pasketti-phonetic achieves an overall phone error rate (PER) of 33% on held-out evaluation data. When compared to leading open phonetic ASR models (like PhoneticXeus) on the same evaluation set, pasketti-phonetic demonstrates improved performance across all demographic groups, cutting the error rate by more than a third for children ages 3-11 and reducing performance disparities between children with and without speech pathologies.
In this model card:
Quickstart
Prerequisites
Hardware. The model can run on CPU or GPU. A GPU is recommended for batches. It was trained and evaluated on Linux with NVIDIA A100 and H100 GPUs.
Environment. The model is a standard
transformersWavLMForCTC:pip install "transformers==4.57.6" "torch==2.6.0" librosaThis release was verified against these versions with Python 3.13.
Data. The model expects 16 kHz mono audio. Each file should hold a single utterance from a single speaker. For best results, clip audio to 60 seconds or less. To match how the model was trained, load the audio with
librosa, as in the example below.
Inference
Pass the Hub repo id, drivendata/pasketti-phonetic. The weights download automatically the first time the model loads.
Transcribe a single clip:
import librosa
from transformers import pipeline
asr = pipeline("automatic-speech-recognition", model="drivendata/pasketti-phonetic")
wav, _ = librosa.load("clip.flac", sr=16000, mono=True)
print(asr(wav)["text"])
>> "ɛn æ dɚæf ɪz deiɹ tu"
Transcribe a batch of clips:
import librosa
import torch
from transformers import Wav2Vec2Processor, WavLMForCTC
repo = "drivendata/pasketti-phonetic"
processor = Wav2Vec2Processor.from_pretrained(repo)
model = WavLMForCTC.from_pretrained(repo).eval()
paths = ["a.flac", "b.flac"]
wavs = [librosa.load(p, sr=16000, mono=True)[0] for p in paths]
inputs = processor(wavs, sampling_rate=16000, padding=True, return_tensors="pt")
with torch.no_grad():
pred_ids = model(**inputs).logits.argmax(-1)
# Blank out frames that only exist because of padding
n_frames = model._get_feat_extract_output_lengths(inputs["attention_mask"].sum(-1))
pred_ids[torch.arange(pred_ids.shape[1]) >= n_frames[:, None]] = processor.tokenizer.pad_token_id
preds = processor.batch_decode(pred_ids)
Usage tips:
- Precision. Load the model in full precision (fp32). Do not pass
torch_dtype=torch.float16ortorch.bfloat16, because WavLM's relative position bias is prone to producing NaNs in half precision. - Decoding. All reported performance uses greedy CTC decoding, as in the snippets above. The model's logits can be passed to an external beam-search decoder such as
pyctcdecode, but that has not been evaluated with this model. - Output. Transcripts use a 55-character IPA alphabet (54 IPA characters plus a character for spaces between words). The alphabet includes combining tilde (nasalization) and the tie bar (diphthongs). Output is not normalized the way it is for evaluation, so normalize both predictions and references the same way before scoring against other transcripts.
- Long audio. The model was trained on clips of at most 30 seconds, so longer audio is out of distribution.
pipelinecan split it into chunks, for exampleasr(wav, chunk_length_s=20, stride_length_s=4).
Model Details
This model is a supervised fine-tune of WavLM Large that transcribes children's speech into IPA phones. During training, the encoder fed two output heads: one predicted IPA phones, and a secondary head predicted words using the subword vocabulary of NVIDIA's Parakeet-TDT 0.6B. The secondary head let the model also learn from utterances with only a word-level transcript, which greatly outnumbered those with phonetic labels. Only the IPA head is used at inference, so the model outputs a sequence of IPA symbols.
The model was fine-tuned for five epochs using AdamW with learning rates: 3e-5 for the pretrained backbone and 1e-4 for the task head. The LR was warmed up linearly over the first 10% of training then decayed to zero on a cosine schedule. Weight decay of 0.01 was applied to all weights except biases and other 1-D parameters (e.g., LayerNorm). Training was completed on a single H100 machine with a batch size of four with two gradient accumulation steps.
The model design is based on the results of the On Top of Pasketti Challenge, in which participants competed to build child ASR models.
Developed by: DrivenData
References
Uses
This model is intended to be used for automatic speech recognition in service of improving educational and developmental outcomes for children.
Based on the model's performance, it is most suitable for downstream use cases that depend on child transcripts as an input. These downstream use cases often do not require exact transcripts to effectively inform decision-making.
For use cases that require exact transcription from child speech, model outputs should be human reviewed.
Downstream Use Cases
Example downstream use cases include:
- Needs screening: Identifying children who may be struggling and better targeting additional support. For example, identifying children behind in literacy or who may have a speech pathology.
- Formative assessment: Monitoring student learning based on naturally collected classroom speech data.
- Educator feedback: Enabling better teacher or tutor support and growth by incorporating spoken interactions.
The model could be plugged into decision-making processes as is to generate transcripts for human review, or fine-tuned to predict in-scope outcomes directly. Intended users include education researchers, education technology developers, teachers, and school administrators.
Out-of-Scope Use
The model should not be used in ways that do not serve the interests of children whose speech it processes. Out-of-scope uses include:
- Surveillance or profiling of children, such as building voice or behavioral profiles for advertising, identifying individual speakers, or monitoring children without informed consent from them and their guardians.
- Engagement maximization, such as designing products that use speech signals to drive compulsive use or unhealthy dependence on technology.
- Fully automated high-stakes decisions, such as diagnosis or educational placement, without review by a qualified educator or clinician.
Evaluation
Model Performance
pasketti-phonetic achieves an overall phone error rate (PER) of 33% on held-out evaluation data, compared with 59% for PhoneticXeus, a leading open phone recognizer. Unlike general-purpose phone recognizers, pasketti-phonetic also marks word boundaries, so its phonetic transcripts line up word by word with what the child was saying. pasketti-phonetic has a 30% PER when counting word boundaries as symbols.
The evaluation set includes 44 hours of child speech from 737 different children as young as 3. All models are scored on every utterance. PhoneticXeus returns no phones for about 15% of utterances, and each of those counts as a full deletion. Before scoring, transcripts are normalized to remove spaces, diphthong tie bars, nasalization combining tildes, stress marks, and ASCII punctuation.
About Phone Error Rate (PER)
Model performance is calculated using IPA Character Error Rate. Character error rate is roughly equivalent to Phone Error Rate (PER) since transcripts are normalized to have roughly one-to-one character-to-phone mapping.
The metric computes the minimum number of substitutions (𝑆), deletions (𝐷), and insertions (𝐼) required to transform the predicted phone sequence into the reference sequence, divided by the total number of reference phones (𝑁). PER indicates roughly what percent of phones predicted are wrong.
A limitation of PER is that it only measures whether the predicted phone is exactly correct — it does not reward similar predictions that are close to correct. As a result, PER often slightly underestimates how well a model can perform in a real-world setting.
Performance by Subset
The model was evaluated for bias in performance based on age, race, speech development pathology, sex, and source corpus. As expected, performance tends to improve for older children compared to younger children. Performance gains are strong across all demographic groups.
Performance by Age
Representation in the evaluation set:
| Age | Utterances | Hours | Children |
|---|---|---|---|
| 3-4 | 50,116 | 23.5 | 189 |
| 5-7 | 10,789 | 10.2 | 293 |
| 8-11 | 8,405 | 3.5 | 114 |
| 12+ | 4,197 | 1.2 | 42 |
| Unknown | 7,626 | 5.3 | 109 |
Performance by Pathology
Representation in the evaluation set:
| Pathology | Utterances | Hours | Children |
|---|---|---|---|
| Atypical development | 31,024 | 11.7 | 168 |
| Typical | 28,789 | 14.7 | 129 |
| Unknown | 21,320 | 17.4 | 440 |
Performance by Race
Racial groups are compared within each individual corpus to remove any correlation between corpus difficulty and race that influences overall race-based performance. This is necessary because very few corpora have race populated.
Representation in the evaluation set:
| Corpus | Race | Utterances | Hours | Children |
|---|---|---|---|---|
| SPROUT | Black | 19,112 | 8.7 | 66 |
| Hispanic | 1,635 | 0.7 | 6 | |
| Multi-racial | 13,278 | 5.3 | 47 | |
| White | 12,452 | 5.7 | 45 | |
| ReadNet | Black | 529 | 0.5 | 19 |
| Multi-racial | 428 | 0.4 | 14 | |
| White | 3,410 | 3.5 | 118 |
Performance by Sex
Representation in the evaluation set:
| Sex | Utterances | Hours | Children |
|---|---|---|---|
| Female | 38,003 | 19.7 | 293 |
| Male | 38,815 | 19.8 | 293 |
| Unknown | 4,315 | 4.2 | 151 |
Performance by Duration
Representation in the evaluation set:
| Utterance duration (seconds) | Utterances | Hours | Children |
|---|---|---|---|
| <1 | 25,592 | 5.5 | 488 |
| [1-3) | 43,265 | 19.3 | 661 |
| [3-5) | 7,160 | 7.6 | 577 |
| [5-10) | 4,366 | 8.2 | 582 |
| [10-20) | 644 | 2.4 | 185 |
| [20-60) | 106 | 0.8 | 57 |
Performance by Corpus
Performance is heavily influenced by corpus due to corpus-specific age ranges, data collection methods, and audio quality. One corpus was put entirely in the test set: JIBO Kids.
Representation in the evaluation set:
| Corpus | Utterances | Hours | Children |
|---|---|---|---|
| Arizona Child Acoustic Database Repository | 703 | 1.2 | 5 |
| Cameron | 1,660 | 1.3 | 5 |
| Edmonton Narrative Norms Instrument | 2,367 | 2.4 | 34 |
| Ellis Weismer Corpus | 1,421 | 0.9 | 5 |
| JIBO Kids† | 7,626 | 5.3 | 109 |
| PERCEPT-GFTA | 4,421 | 1 | 77 |
| PERCEPT-R | 7,970 | 2.6 | 29 |
| ReadNet | 8,230 | 8.3 | 307 |
| Speech Production Repository for Optimizing Use of AI Technologies (SPROUT) | 46,735 | 20.6 | 166 |
† Present only in the evaluation set and not in the training set.
Bias and Limitations
The model performs best on audio data that is:
- Already diarized and clipped to a single utterance. The model was trained on pre-diarized audio clips of individual utterances. It was not trained to perform forced alignment or speaker identification.
- Under 60 seconds in length. The model was trained on utterances that were 30 seconds in length or shorter, and evaluated on audio clips that were less than 60 seconds long. To perform inference on longer audio, we recommend splitting it into chunks first.
The model performs better in some circumstances, and on some individuals, than others. The model struggles more with:
- Very young children (under 5). As expected, the model performs better on older children.
- Children with speech pathologies. Model performance is worse for children with atypical speech development than for typically developing children. Within every corpus that includes both groups, atypical children have a higher error rate, but the size of the gap varies widely by corpus, from a few points in SPROUT and Weismer to much larger gaps in ENNI and PERCEPT-R.
The model has not been evaluated on classroom audio. The only classroom audio in training was orthographically transcribed and used by the secondary word-prediction head. None of the phonetically transcribed training or evaluation data was recorded in a classroom. Classroom audio roughly triples pasketti-word's error rate, so phonetic performance in classrooms is likely worse than reported here.
Performance could not be evaluated for all racial groups. For example, there was insufficient data for Asian children to determine model performance. Within the model's evaluation dataset, performance is broadly similar across the racial groups with enough data to compare, though not identical. Within SPROUT, error rates are somewhat higher for multi-racial and Black children than for White and Hispanic children. Within ReadNet, error rates are higher for Black children than for White children.
Training Data
The model was trained on 116 hours of read, prompted, and spontaneous child speech spanning roughly 1,700 different children as young as 3. An additional 371 hours of orthographically transcribed child speech were used to fine-tune a secondary word prediction head (see the Model Details section).
Training data was intentionally gathered to target specific populations where performance of currently available models is poor:
- Children with atypical speech development
- Non-white children
- Children from low-income backgrounds
Training Data Collection
Training data was compiled as part of the On Top of Pasketti Challenge. Phonetic training data comes from 8 different source corpora (one additional corpus is used as evaluation data only).
Only a few corpora have pre-existing transcriptions. To supplement existing transcriptions, DrivenData managed a team of linguists to transcribe additional audio data. 10% of audio files were retranscribed for quality assurance.
Child speech data is sensitive and difficult to share. We are grateful for the work of the data providers who thoughtfully collected and provided the data used in this project under appropriate consents, and to the speakers represented whose voices have helped to advance this work.
Representation in the Training Data
The breakdowns below describe phonetically transcribed training data. They do not include the additional orthographically transcribed data that was used to train a secondary word-prediction head.
Breakdown by age:
| Age | Utterances | Hours | Children |
|---|---|---|---|
| 3-4 | 52,206 | 34.5 | 241 |
| 5-7 | 43,012 | 42.2 | 828 |
| 8-11 | 75,651 | 30.7 | 504 |
| 12+ | 30,904 | 8.2 | 199 |
| Unknown | 114 | 0.1 | 1 |
Breakdown by pathology:
| Pathology | Utterances | Hours | Children |
|---|---|---|---|
| Atypical development | 122,875 | 45.1 | 553 |
| Typical | 41,496 | 32 | 400 |
| Unknown | 37,516 | 38.6 | 745 |
Breakdown by race:
| Race | Utterances | Hours | Children |
|---|---|---|---|
| Asian | 384 | 0.4 | 15 |
| Black | 7,142 | 3.9 | 84 |
| Hispanic | 606 | 0.3 | 1 |
| Multi-racial | 6,637 | 3.8 | 55 |
| Native American | 11 | <0.1 | 1 |
| Other | 177 | 0.2 | 10 |
| Pacific Islander | 49 | <0.1 | 2 |
| Unknown | 175,996 | 99.1 | 1,298 |
| White | 10,885 | 8.0 | 231 |
Breakdown by utterance duration:
| Utterance duration (seconds) | Utterances | Hours | Children |
|---|---|---|---|
| <1 | 75,640 | 15.7 | 1,066 |
| [1-3) | 88,704 | 39.1 | 1,478 |
| [3-5) | 19,763 | 21.2 | 1,118 |
| [5-10) | 14,625 | 27.6 | 1,142 |
| [10-20) | 2,984 | 10.9 | 395 |
| [20-60) | 171 | 1.2 | 75 |
Breakdown by corpus:
| Corpus | Utterances | Hours | Children |
|---|---|---|---|
| Arizona Child Acoustic Database Repository | 8,390 | 14.5 | 46 |
| Cameron | 9,224 | 6 | 35 |
| Edmonton Narrative Norms Instrument | 21,541 | 25.3 | 311 |
| Ellis Weismer Corpus | 23,389 | 15.6 | 86 |
| PERCEPT-GFTA | 14,996 | 3.2 | 264 |
| PERCEPT-R | 91,071 | 27 | 257 |
| ReadNet | 16,658 | 16.8 | 648 |
| Speech Production Repository for Optimizing Use of AI Technologies (SPROUT) | 16,618 | 7.4 | 50 |
Data Processing
Data was prepared for training by:
- Separating into utterances. Audio was segmented to create a single audio file per utterance clipped to the utterance boundaries. Each sample was a single utterance audio with a corresponding transcript.
- Normalizing transcriptions. Transcripts were normalized to standardize IPA symbols and map each phone to roughly one character. Nasalization (combining tildes) and diphthong tie bars (e.g. e͡ɪ) were preserved.
- Resampling. Audio data was resampled to 16 kHz, 1-channel, 16-bit signed and converted to FLAC format.
- Long utterance clipping. Clips longer than 30 seconds were cropped to a random 30-second window.
- Concatenating. During training, up to half of the samples had other phonetically transcribed utterances appended, up to 30 seconds total, with their transcripts joined. Most utterances in the training set are only one to three words long, and this exposes the model to longer, multi-phrase audio.
Have questions or want to know more? Get in touch!
Citation
@misc{drivendata2026pasketti,
author = {{DrivenData}},
title = {{pasketti-phonetic}: A speech recognition model for children},
year = {2026},
publisher = {Hugging Face},
howpublished = {Machine learning model},
url = {https://huggingface.co/drivendata/pasketti-phonetic}
}
We are grateful to the many individuals, organizations, and teams who contributed expertise, data, and guidance throughout the development of this project. For full project acknowledgements, see the On Top of Pasketti Challenge page.
- Downloads last month
- 4
Model tree for drivendata/pasketti-phonetic
Base model
microsoft/wavlm-large




