SpeakClear AI

SpeakClear AI is a non-clinical pronunciation-practice research project. It uses acoustic features and machine learning to classify short pronunciation recordings as clear or unclear.

I built this project because I wanted to explore whether a small personalized speech dataset could still support useful pronunciation-practice feedback. The goal is not to diagnose or treat speech issues, but to create a simple educational tool for practice and research.

Live Demo

Hugging Face demo: https://huggingface.co/spaces/rvarsh/SpeakClear-AI-Demo

Streamlit demo: https://speakclear-ai.streamlit.app/

GitHub

https://github.com/rohilvarshney/SpeakClearAI

Intended Use

SpeakClear AI is meant for:

  • educational pronunciation practice
  • personal speech-practice feedback
  • machine-learning research with small speech datasets
  • exploring how public pronunciation data transfers to personalized recordings

It is not meant for diagnosis, treatment, grading, or clinical decision-making.

Dataset Summary

The personalized SpeakClear dataset contains 242 usable labeled clips after filtering for available audio files.

Target categories:

  • /s/
  • /z/
  • /r/
  • /th/
  • mixed /s/ + /z/ sentence contexts

Raw personal voice recordings are not publicly released.

Public Dataset Experiment

I also tested the same feature pipeline on target-phone examples extracted from the public SpeechOcean762 pronunciation dataset.

The project compared:

  1. SpeakClear-only training
  2. SpeechOcean-only training
  3. SpeechOcean + SpeakClear transfer training

Main SpeakClear Results

Across 30 repeated stratified splits of the SpeakClear dataset:

Model Accuracy Balanced Accuracy Unclear Recall Unclear F1
SVM 0.693 ± 0.050 0.671 ± 0.062 0.629 ± 0.133 0.483 ± 0.074
Random Forest 0.760 ± 0.043 0.640 ± 0.064 0.419 ± 0.128 0.440 ± 0.106
Logistic Regression 0.627 ± 0.048 0.591 ± 0.061 0.524 ± 0.130 0.389 ± 0.073
Majority Baseline 0.770 ± 0.000 0.500 ± 0.000 0.000 ± 0.000 0.000 ± 0.000

The majority baseline reached high accuracy because the dataset is imbalanced, but it completely failed to detect unclear clips. This shows why accuracy alone is misleading.

Transfer Result

The transfer experiment showed that public pronunciation data helped some models, especially Random Forest, but it did not fully replace personalized data.

On the held-out SpeakClear test set:

  • SpeakClear-only SVM unclear F1: 0.579
  • SpeechOcean + SpeakClear Random Forest unclear F1: 0.571
  • SpeechOcean + SpeakClear Random Forest accuracy: 0.803

This suggests that public pronunciation data can improve robustness, but personalized recordings still matter because of domain mismatch.

Included Files

This repository includes:

  • models/speakclear_random_forest.joblib
  • SpeakClear evaluation results
  • SpeechOcean evaluation results
  • transfer experiment results

It does not include raw voice recordings.

Limitations

  • single-speaker personalized SpeakClear dataset
  • small number of unclear clips
  • manually assigned labels
  • not clinically validated
  • may not generalize to other speakers
  • public-dataset transfer is limited by domain mismatch
  • not a replacement for a speech-language pathologist

Safety Notice

SpeakClear AI is a research prototype and educational speech-practice tool. It is not a diagnosis tool, treatment tool, or replacement for a speech-language pathologist.

Downloads last month
-
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support