Research Project Overview and Description
The computational language learning simulations conducted by the Language Development Lab at the Department of Linguistics collectively contribute to developing a unified computational framework for early phonological acquisition, demonstrating through unsupervised neural network models—such as autoencoders, LSTMs, CNNs, ResNets, and attention-based architectures—how bottom-up learning from raw acoustic input gives rise to universal phonetic sensitivities as a transient foundational stage before transitioning to language-specific perceptual attunement.
The research reveals the emergence of phoneme-like categories, feature-aligned representations, prosodic sensitivities (e.g., stress and tone), and phonotactic regularities, with models exhibiting mosaic, feature-dependent trajectories that mirror infants’ perceptual narrowing in the first year of life.
A key overarching insight is the facilitative role of staged sensory input: initial exposure to degraded (low-pass filtered) signals promotes robust hidden representations, accelerated convergence on contrasts, superior generalization, and efficient bootstrapping upon encountering full-spectrum speech, outperforming unfiltered or mismatched conditions while underscoring vulnerabilities associated with atypical early auditory environments.
Research Outcome
This mechanistic account highlights how hierarchical, data-driven processes enable the shift from broad universal discrimination to native-language specialization, providing explanatory power for typical and atypical developmental pathways in phonetic, prosodic, and phonological learning.
About the Researcher
Dr. Youngah Do is an Associate Professor in the Department of Linguistics at HKU, holding a PhD from MIT. Her research examines language learning and learnability through experimental and computational modeling, with recent work focusing on the preservation and inclusivity of Hong Kong Sign Language.
Publications
Tan, F. L., & Do, Y. (2025a). Attention-LSTM autoencoder simulation for phonotactic learning from raw audio input. Linguistics Vanguard. https://doi.org/10.1515/lingvan-2024-0210
Tan, F. L., & Do, Y. (2025b). Bottom-up modeling of phoneme learning: Universal sensitivity and language-specific transformation. Speech Communication, 176, 103343. https://doi.org/10.1016/j.specom.2025.103343
Tan, F. L., Zheng, S., Liu, M., & Do, Y. (2026). Modeling Prosodic Development with Prenatal Audio Attenuation. University of Massachusetts (UMass) Amherst. https://doi.org/10.7275/amphonology.3665
Zheng, S., Tan, F. L., & Do, Y. (2026). Modeling the impact of prenatal audio attenuation on speech sound learning. Journal of Experimental Psychology Learning Memory and Cognition. https://doi.org/10.1037/xlm0001616
Fund Source
N/A
For inquiries, please contact
atlab@hku.hk
For more information, please refer to the website
Language Development Lab | Department of Linguistics | HKU
