pasketti-phonetic 🍝

pasketti-phonetic is an automatic speech recognition (ASR) model designed specifically for child speech. pasketti-phonetic generates phonetic (per-sound) transcriptions in the International Phonetic Alphabet (IPA), including nasalization and diphthong markings, from audio files.

Audio in, phones out

Phonetic transcript of a child saying "and a giraffe is there too." The model predicts both spoken phones and word boundaries.

🐣 Existing open-source ASR models trained on adult speech perform poorly on children because child speech is fundamentally different from adult speech. Children are still learning to produce speech sounds and developing fine motor skills. Common child speech errors include:

  • Metathesis: "elephant" → "ephelant"
  • Velar fronting: "cup" → "tup"
  • Syllable deletion: "banana" → "nana"
  • Combinations of errors: "spaghetti" → "pasketti" 🍝

🔉 pasketti-phonetic is fine-tuned on a large corpus of 116 hours of phonetically transcribed child speech spanning roughly 1,700 different children. This is the largest dataset of its kind for use in developing and testing open ASR models for kids. The model is designed to advance early education assessments, teaching tools, and needs screening.

📈 pasketti-phonetic achieves an overall phone error rate (PER) of 33% on held-out evaluation data. When compared to leading open phonetic ASR models (like PhoneticXeus) on the same evaluation set, pasketti-phonetic demonstrates improved performance across all demographic groups, cutting the error rate by more than a third for children ages 3-11 and reducing performance disparities between children with and without speech pathologies.

In this model card:

  1. Quickstart
  2. Model Details
  3. Uses
  4. Evaluation
    1. Performance by Subset
    2. Bias and Limitations
  5. Training Data

Quickstart

Prerequisites

  1. Hardware. The model can run on CPU or GPU. A GPU is recommended for batches. It was trained and evaluated on Linux with NVIDIA A100 and H100 GPUs.

  2. Environment. The model is a standard transformers WavLMForCTC:

    pip install "transformers==4.57.6" "torch==2.6.0" librosa
    

    This release was verified against these versions with Python 3.13.

  3. Data. The model expects 16 kHz mono audio. Each file should hold a single utterance from a single speaker. For best results, clip audio to 60 seconds or less. To match how the model was trained, load the audio with librosa, as in the example below.

Inference

Pass the Hub repo id, drivendata/pasketti-phonetic. The weights download automatically the first time the model loads.

Transcribe a single clip:

import librosa
from transformers import pipeline

asr = pipeline("automatic-speech-recognition", model="drivendata/pasketti-phonetic")
wav, _ = librosa.load("clip.flac", sr=16000, mono=True)
print(asr(wav)["text"])

>> "ɛn æ dɚæf ɪz deiɹ tu"

Transcribe a batch of clips:

import librosa
import torch
from transformers import Wav2Vec2Processor, WavLMForCTC

repo = "drivendata/pasketti-phonetic"
processor = Wav2Vec2Processor.from_pretrained(repo)
model = WavLMForCTC.from_pretrained(repo).eval()

paths = ["a.flac", "b.flac"]
wavs = [librosa.load(p, sr=16000, mono=True)[0] for p in paths]
inputs = processor(wavs, sampling_rate=16000, padding=True, return_tensors="pt")

with torch.no_grad():
    pred_ids = model(**inputs).logits.argmax(-1)

# Blank out frames that only exist because of padding
n_frames = model._get_feat_extract_output_lengths(inputs["attention_mask"].sum(-1))
pred_ids[torch.arange(pred_ids.shape[1]) >= n_frames[:, None]] = processor.tokenizer.pad_token_id
preds = processor.batch_decode(pred_ids)

Usage tips:

  • Precision. Load the model in full precision (fp32). Do not pass torch_dtype=torch.float16 or torch.bfloat16, because WavLM's relative position bias is prone to producing NaNs in half precision.
  • Decoding. All reported performance uses greedy CTC decoding, as in the snippets above. The model's logits can be passed to an external beam-search decoder such as pyctcdecode, but that has not been evaluated with this model.
  • Output. Transcripts use a 55-character IPA alphabet (54 IPA characters plus a character for spaces between words). The alphabet includes combining tilde (nasalization) and the tie bar (diphthongs). Output is not normalized the way it is for evaluation, so normalize both predictions and references the same way before scoring against other transcripts.
  • Long audio. The model was trained on clips of at most 30 seconds, so longer audio is out of distribution. pipeline can split it into chunks, for example asr(wav, chunk_length_s=20, stride_length_s=4).

Model Details

This model is a supervised fine-tune of WavLM Large that transcribes children's speech into IPA phones. During training, the encoder fed two output heads: one predicted IPA phones, and a secondary head predicted words using the subword vocabulary of NVIDIA's Parakeet-TDT 0.6B. The secondary head let the model also learn from utterances with only a word-level transcript, which greatly outnumbered those with phonetic labels. Only the IPA head is used at inference, so the model outputs a sequence of IPA symbols.

The model was fine-tuned for five epochs using AdamW with learning rates: 3e-5 for the pretrained backbone and 1e-4 for the task head. The LR was warmed up linearly over the first 10% of training then decayed to zero on a cosine schedule. Weight decay of 0.01 was applied to all weights except biases and other 1-D parameters (e.g., LayerNorm). Training was completed on a single H100 machine with a batch size of four with two gradient accumulation steps.

The model design is based on the results of the On Top of Pasketti Challenge, in which participants competed to build child ASR models.

Uses

This model is intended to be used for automatic speech recognition in service of improving educational and developmental outcomes for children.

Based on the model's performance, it is most suitable for downstream use cases that depend on child transcripts as an input. These downstream use cases often do not require exact transcripts to effectively inform decision-making.

For use cases that require exact transcription from child speech, model outputs should be human reviewed.

Downstream Use Cases

Example downstream use cases include:

  • Needs screening: Identifying children who may be struggling and better targeting additional support. For example, identifying children behind in literacy or who may have a speech pathology.
  • Formative assessment: Monitoring student learning based on naturally collected classroom speech data.
  • Educator feedback: Enabling better teacher or tutor support and growth by incorporating spoken interactions.

The model could be plugged into decision-making processes as is to generate transcripts for human review, or fine-tuned to predict in-scope outcomes directly. Intended users include education researchers, education technology developers, teachers, and school administrators.

Out-of-Scope Use

The model should not be used in ways that do not serve the interests of children whose speech it processes. Out-of-scope uses include:

  • Surveillance or profiling of children, such as building voice or behavioral profiles for advertising, identifying individual speakers, or monitoring children without informed consent from them and their guardians.
  • Engagement maximization, such as designing products that use speech signals to drive compulsive use or unhealthy dependence on technology.
  • Fully automated high-stakes decisions, such as diagnosis or educational placement, without review by a qualified educator or clinician.

Evaluation

Model Performance

pasketti-phonetic achieves an overall phone error rate (PER) of 33% on held-out evaluation data, compared with 59% for PhoneticXeus, a leading open phone recognizer. Unlike general-purpose phone recognizers, pasketti-phonetic also marks word boundaries, so its phonetic transcripts line up word by word with what the child was saying. pasketti-phonetic has a 30% PER when counting word boundaries as symbols.

PER by model

The evaluation set includes 44 hours of child speech from 737 different children as young as 3. All models are scored on every utterance. PhoneticXeus returns no phones for about 15% of utterances, and each of those counts as a full deletion. Before scoring, transcripts are normalized to remove spaces, diphthong tie bars, nasalization combining tildes, stress marks, and ASCII punctuation.

About Phone Error Rate (PER)

Model performance is calculated using IPA Character Error Rate. Character error rate is roughly equivalent to Phone Error Rate (PER) since transcripts are normalized to have roughly one-to-one character-to-phone mapping.

The metric computes the minimum number of substitutions (𝑆), deletions (𝐷), and insertions (𝐼) required to transform the predicted phone sequence into the reference sequence, divided by the total number of reference phones (𝑁). PER indicates roughly what percent of phones predicted are wrong.

PER=Substitutions+Deletions+InsertionsPhones=S+D+IN \text{PER} = \frac{\text{Substitutions} + \text{Deletions} + \text{Insertions}}{\text{Phones}} = \frac{S + D + I}{N}

A limitation of PER is that it only measures whether the predicted phone is exactly correct — it does not reward similar predictions that are close to correct. As a result, PER often slightly underestimates how well a model can perform in a real-world setting.

Performance by Subset

The model was evaluated for bias in performance based on age, race, speech development pathology, sex, and source corpus. As expected, performance tends to improve for older children compared to younger children. Performance gains are strong across all demographic groups.

Performance by Age

PER by age

Representation in the evaluation set:

Age Utterances Hours Children
3-4 50,116 23.5 189
5-7 10,789 10.2 293
8-11 8,405 3.5 114
12+ 4,197 1.2 42
Unknown 7,626 5.3 109

Performance by Pathology

PER by speech pathology

Representation in the evaluation set:

Pathology Utterances Hours Children
Atypical development 31,024 11.7 168
Typical 28,789 14.7 129
Unknown 21,320 17.4 440

Performance by Race

Racial groups are compared within each individual corpus to remove any correlation between corpus difficulty and race that influences overall race-based performance. This is necessary because very few corpora have race populated.

PER by race and corpus

Representation in the evaluation set:

CorpusRaceUtterancesHoursChildren
SPROUTBlack19,1128.766
Hispanic1,6350.76
Multi-racial13,2785.347
White12,4525.745
ReadNetBlack5290.519
Multi-racial4280.414
White3,4103.5118

Performance by Sex

PER by sex

Representation in the evaluation set:

Sex Utterances Hours Children
Female 38,003 19.7 293
Male 38,815 19.8 293
Unknown 4,315 4.2 151

Performance by Duration

PER by utterance duration

Representation in the evaluation set:

Utterance duration (seconds) Utterances Hours Children
<1 25,592 5.5 488
[1-3) 43,265 19.3 661
[3-5) 7,160 7.6 577
[5-10) 4,366 8.2 582
[10-20) 644 2.4 185
[20-60) 106 0.8 57

Performance by Corpus

Performance is heavily influenced by corpus due to corpus-specific age ranges, data collection methods, and audio quality. One corpus was put entirely in the test set: JIBO Kids.

PER by corpus

Representation in the evaluation set:

Corpus Utterances Hours Children
Arizona Child Acoustic Database Repository 703 1.2 5
Cameron 1,660 1.3 5
Edmonton Narrative Norms Instrument 2,367 2.4 34
Ellis Weismer Corpus 1,421 0.9 5
JIBO Kids† 7,626 5.3 109
PERCEPT-GFTA 4,421 1 77
PERCEPT-R 7,970 2.6 29
ReadNet 8,230 8.3 307
Speech Production Repository for Optimizing Use of AI Technologies (SPROUT) 46,735 20.6 166

† Present only in the evaluation set and not in the training set.

Bias and Limitations

The model performs best on audio data that is:

  • Already diarized and clipped to a single utterance. The model was trained on pre-diarized audio clips of individual utterances. It was not trained to perform forced alignment or speaker identification.
  • Under 60 seconds in length. The model was trained on utterances that were 30 seconds in length or shorter, and evaluated on audio clips that were less than 60 seconds long. To perform inference on longer audio, we recommend splitting it into chunks first.

The model performs better in some circumstances, and on some individuals, than others. The model struggles more with:

  • Very young children (under 5). As expected, the model performs better on older children.
  • Children with speech pathologies. Model performance is worse for children with atypical speech development than for typically developing children. Within every corpus that includes both groups, atypical children have a higher error rate, but the size of the gap varies widely by corpus, from a few points in SPROUT and Weismer to much larger gaps in ENNI and PERCEPT-R.

The model has not been evaluated on classroom audio. The only classroom audio in training was orthographically transcribed and used by the secondary word-prediction head. None of the phonetically transcribed training or evaluation data was recorded in a classroom. Classroom audio roughly triples pasketti-word's error rate, so phonetic performance in classrooms is likely worse than reported here.

Performance could not be evaluated for all racial groups. For example, there was insufficient data for Asian children to determine model performance. Within the model's evaluation dataset, performance is broadly similar across the racial groups with enough data to compare, though not identical. Within SPROUT, error rates are somewhat higher for multi-racial and Black children than for White and Hispanic children. Within ReadNet, error rates are higher for Black children than for White children.


Training Data

The model was trained on 116 hours of read, prompted, and spontaneous child speech spanning roughly 1,700 different children as young as 3. An additional 371 hours of orthographically transcribed child speech were used to fine-tune a secondary word prediction head (see the Model Details section).

Training data was intentionally gathered to target specific populations where performance of currently available models is poor:

  • Children with atypical speech development
  • Non-white children
  • Children from low-income backgrounds

Training Data Collection

Training data was compiled as part of the On Top of Pasketti Challenge. Phonetic training data comes from 8 different source corpora (one additional corpus is used as evaluation data only).

Only a few corpora have pre-existing transcriptions. To supplement existing transcriptions, DrivenData managed a team of linguists to transcribe additional audio data. 10% of audio files were retranscribed for quality assurance.

Child speech data is sensitive and difficult to share. We are grateful for the work of the data providers who thoughtfully collected and provided the data used in this project under appropriate consents, and to the speakers represented whose voices have helped to advance this work.

Representation in the Training Data

The breakdowns below describe phonetically transcribed training data. They do not include the additional orthographically transcribed data that was used to train a secondary word-prediction head.

Breakdown by age:

Age Utterances Hours Children
3-4 52,206 34.5 241
5-7 43,012 42.2 828
8-11 75,651 30.7 504
12+ 30,904 8.2 199
Unknown 114 0.1 1

Breakdown by pathology:

Pathology Utterances Hours Children
Atypical development 122,875 45.1 553
Typical 41,496 32 400
Unknown 37,516 38.6 745

Breakdown by race:

Race Utterances Hours Children
Asian 384 0.4 15
Black 7,142 3.9 84
Hispanic 606 0.3 1
Multi-racial 6,637 3.8 55
Native American 11 <0.1 1
Other 177 0.2 10
Pacific Islander 49 <0.1 2
Unknown 175,996 99.1 1,298
White 10,885 8.0 231

Breakdown by utterance duration:

Utterance duration (seconds) Utterances Hours Children
<1 75,640 15.7 1,066
[1-3) 88,704 39.1 1,478
[3-5) 19,763 21.2 1,118
[5-10) 14,625 27.6 1,142
[10-20) 2,984 10.9 395
[20-60) 171 1.2 75

Breakdown by corpus:

Corpus Utterances Hours Children
Arizona Child Acoustic Database Repository 8,390 14.5 46
Cameron 9,224 6 35
Edmonton Narrative Norms Instrument 21,541 25.3 311
Ellis Weismer Corpus 23,389 15.6 86
PERCEPT-GFTA 14,996 3.2 264
PERCEPT-R 91,071 27 257
ReadNet 16,658 16.8 648
Speech Production Repository for Optimizing Use of AI Technologies (SPROUT) 16,618 7.4 50

Data Processing

Data was prepared for training by:

  • Separating into utterances. Audio was segmented to create a single audio file per utterance clipped to the utterance boundaries. Each sample was a single utterance audio with a corresponding transcript.
  • Normalizing transcriptions. Transcripts were normalized to standardize IPA symbols and map each phone to roughly one character. Nasalization (combining tildes) and diphthong tie bars (e.g. e͡ɪ) were preserved.
  • Resampling. Audio data was resampled to 16 kHz, 1-channel, 16-bit signed and converted to FLAC format.
  • Long utterance clipping. Clips longer than 30 seconds were cropped to a random 30-second window.
  • Concatenating. During training, up to half of the samples had other phonetically transcribed utterances appended, up to 30 seconds total, with their transcripts joined. Most utterances in the training set are only one to three words long, and this exposes the model to longer, multi-phrase audio.

Have questions or want to know more? Get in touch!


Citation

@misc{drivendata2026pasketti,
  author       = {{DrivenData}},
  title        = {{pasketti-phonetic}: A speech recognition model for children},
  year         = {2026},
  publisher    = {Hugging Face},
  howpublished = {Machine learning model},
  url          = {https://huggingface.co/drivendata/pasketti-phonetic}
}

We are grateful to the many individuals, organizations, and teams who contributed expertise, data, and guidance throughout the development of this project. For full project acknowledgements, see the On Top of Pasketti Challenge page.

Downloads last month
4
Safetensors
Model size
0.3B params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for drivendata/pasketti-phonetic

Finetuned
(37)
this model

Collection including drivendata/pasketti-phonetic