Speakers
Description
This presentation addresses the problem of creating accurate spoken L2 learner corpora by introducing a hybrid transcription approach combining top-down LLM-mediated ASR (Whisper) with bottom-up wav2vec. Findings from an analysis of word error rate in 200 L2 speech samples show that ASR accuracy varies by task and proficiency. The presenters will demonstrate a prototype tool that aligns ASR and wav2vec transcripts to enable efficient human review and to export a learner corpus for further analysis.
Abstract section 3: Content/method
The presenters report findings from an analysis of word error rate (WER) in 200 samples of L2-generated speech, including 100 presentations and 100 discussions from learners at CEFR A2 to B2. ASR transcripts were combined with wav2vec output using a modified Needleman-Wunsch alignment algorithm. The resulting transcript was compared against professionally transcribed and manually cleaned versions. These findings were then used to identify problematic phonemes in L2 learner speech to further improve the wav2vec models.
Abstract section 5: References
Baevski, A., Zhou, Y., Mohamed, A., & Auli, M. (2020). wav2vec 2.0: A framework for self-supervised learning of speech representations. Advances in Neural Information Processing Systems, 33, 12449–12460.
Brezina, V., Gablasova, D., & Lenko-Szymanska, A. (2022). Learner corpus research: Past, present and future. In V. Brezina, D. Gablasova, & A. Lenko-Szymanska (Eds.), The Routledge handbook of corpora and English language teaching and learning (pp. 107–123). Routledge.
Kyle, K., & Eguchi, M. (2024). Assessing the accuracy of automatic speech recognition for L2 English. Language Learning & Technology, 28(1), 1–23.
McGuire, M., & Larson-Hall, J. (2025). Assessing Whisper automatic speech recognition and WER scoring for elicited imitation: Steps toward automation. Research Methods in Applied Linguistics, 4(1), 100197. https://doi.org/10.1016/j.rmal.2025.100197
Radford, A., Kim, J. W., Xu, T., Brockman, G., McLeavey, C., & Sutskever, I. (2022). Robust speech recognition via large-scale weak supervision. arXiv preprint arXiv:2212.04356.
Wang, Y., Shen, J., & Jia, Y. (2021). Non-native English speaker speech recognition challenges. Proceedings of Interspeech 2021, 3820–3824.
Abstract section 4: Outcomes/results
Results indicate that while ASR performs comparably to professional transcription for single-speaker presentations, it produces a significantly higher WER for multi-speaker discussions due to overlapping speech, varying accents, and background noise. Higher-proficiency participants' speech was transcribed with greater accuracy. Based on these findings and insights from prior literature, the presenter introduces a prototype tool that aligns ASR and word2vec transcripts at the word level, highlights discrepancies for manual review, and exports time-stamped gold-standard transcripts suitable for corpus analysis in tools such as PRAAT. We will also provide recommendations for recording and using ASR in L2 research contexts.
Abstract section 2: Contribution/research questions
This presentation investigates a hybrid transcription approach that combines top-down LLM-mediated ASR (Whisper) with bottom-up wav2vec for learner corpus development. Drawing on the findings from previous projects that have examined ASR accuracy in L2 speech (McGuire & Larson-Hall, 2025), the presenters investigate two research questions:
1. How accurate are these approaches compared to a human transcriber in capturing vocabulary used by learners?
2. Can the system's performance be improved by training wav2vec on learner-generated speech?
Abstract section 1: Relevance
Recent studies have explored the use of Automated Speech Recognition (ASR) models like Whisper for transcribing L1 speech (Baevski et al., 2020; Radford et al., 2022), but there are challenges with applying ASR to L2 English learners due to pronunciation errors, disfluencies, and atypical grammatical constructions (Kyle & Eguchi, 2024; Wang et al., 2021). Because of this, there is a need for better tools that integrate ASR output with human oversight to create reliable transcripts, as a means of building spoken corpora to facilitate research on learner vocabulary usage across different domains and language contexts (Brezina et al., 2022).
| Title | Improving Automated Transcription of L2 Learner Texts |
|---|---|
| Teaching Context | College and university education |