TL;DR
Korean pediatric speech disorders lack effective automated assessment tools. An end-to-end pipeline was developed using neural speaker diarization and self-supervised learning for pronunciation evaluation.
✦ Why It Matters
Engineers can leverage this automated evaluation pipeline to improve speech assessment tools for young children.
Key Takeaways
How It Works
The system employs neural speaker diarization to accurately identify and separate speech from toddlers and their caregivers. This is crucial for evaluating toddler pronunciation without interference from adult speech patterns.
The self-supervised learning models, specifically HuBERT-large and WavLM-large, are used to assess the correctness of consonants and vowels, respectively, enhancing the evaluation process.
Related