Instructions to use rvarsh/SpeakClear-AI with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Scikit-learn
How to use rvarsh/SpeakClear-AI with Scikit-learn:
from huggingface_hub import hf_hub_download import joblib model = joblib.load( hf_hub_download("rvarsh/SpeakClear-AI", "sklearn_model.joblib") ) # only load pickle files from sources you trust # read more about it here https://skops.readthedocs.io/en/stable/persistence.html - Notebooks
- Google Colab
- Kaggle
SpeakClear AI
SpeakClear AI is a non-clinical pronunciation-practice research project. It uses acoustic features and machine learning to classify short pronunciation recordings as clear or unclear.
I built this project because I wanted to explore whether a small personalized speech dataset could still support useful pronunciation-practice feedback. The goal is not to diagnose or treat speech issues, but to create a simple educational tool for practice and research.
Live Demo
Hugging Face demo: https://huggingface.co/spaces/rvarsh/SpeakClear-AI-Demo
Streamlit demo: https://speakclear-ai.streamlit.app/
GitHub
https://github.com/rohilvarshney/SpeakClearAI
Intended Use
SpeakClear AI is meant for:
- educational pronunciation practice
- personal speech-practice feedback
- machine-learning research with small speech datasets
- exploring how public pronunciation data transfers to personalized recordings
It is not meant for diagnosis, treatment, grading, or clinical decision-making.
Dataset Summary
The personalized SpeakClear dataset contains 242 usable labeled clips after filtering for available audio files.
Target categories:
- /s/
- /z/
- /r/
- /th/
- mixed /s/ + /z/ sentence contexts
Raw personal voice recordings are not publicly released.
Public Dataset Experiment
I also tested the same feature pipeline on target-phone examples extracted from the public SpeechOcean762 pronunciation dataset.
The project compared:
- SpeakClear-only training
- SpeechOcean-only training
- SpeechOcean + SpeakClear transfer training
Main SpeakClear Results
Across 30 repeated stratified splits of the SpeakClear dataset:
| Model | Accuracy | Balanced Accuracy | Unclear Recall | Unclear F1 |
|---|---|---|---|---|
| SVM | 0.693 ± 0.050 | 0.671 ± 0.062 | 0.629 ± 0.133 | 0.483 ± 0.074 |
| Random Forest | 0.760 ± 0.043 | 0.640 ± 0.064 | 0.419 ± 0.128 | 0.440 ± 0.106 |
| Logistic Regression | 0.627 ± 0.048 | 0.591 ± 0.061 | 0.524 ± 0.130 | 0.389 ± 0.073 |
| Majority Baseline | 0.770 ± 0.000 | 0.500 ± 0.000 | 0.000 ± 0.000 | 0.000 ± 0.000 |
The majority baseline reached high accuracy because the dataset is imbalanced, but it completely failed to detect unclear clips. This shows why accuracy alone is misleading.
Transfer Result
The transfer experiment showed that public pronunciation data helped some models, especially Random Forest, but it did not fully replace personalized data.
On the held-out SpeakClear test set:
- SpeakClear-only SVM unclear F1: 0.579
- SpeechOcean + SpeakClear Random Forest unclear F1: 0.571
- SpeechOcean + SpeakClear Random Forest accuracy: 0.803
This suggests that public pronunciation data can improve robustness, but personalized recordings still matter because of domain mismatch.
Included Files
This repository includes:
models/speakclear_random_forest.joblib- SpeakClear evaluation results
- SpeechOcean evaluation results
- transfer experiment results
It does not include raw voice recordings.
Limitations
- single-speaker personalized SpeakClear dataset
- small number of unclear clips
- manually assigned labels
- not clinically validated
- may not generalize to other speakers
- public-dataset transfer is limited by domain mismatch
- not a replacement for a speech-language pathologist
Safety Notice
SpeakClear AI is a research prototype and educational speech-practice tool. It is not a diagnosis tool, treatment tool, or replacement for a speech-language pathologist.
- Downloads last month
- -