Get our free extension to see links to code for papers anywhere online!Free add-on: code for papers everywhere!Free add-on: See code for papers anywhere!

Add to Chrome

Add to Firefox

Add to Edge

William Lamb

A Practitioner's Guide to Building ASR Models for Low-Resource Languages: A Case Study on Scottish Gaelic

Jun 05, 2025

Ondřej Klejch, William Lamb, Peter Bell

Abstract:An effective approach to the development of ASR systems for low-resource languages is to fine-tune an existing multilingual end-to-end model. When the original model has been trained on large quantities of data from many languages, fine-tuning can be effective with limited training data, even when the language in question was not present in the original training data. The fine-tuning approach has been encouraged by the availability of public-domain E2E models and is widely believed to lead to state-of-the-art results. This paper, however, challenges that belief. We show that an approach combining hybrid HMMs with self-supervised models can yield substantially better performance with limited training data. This combination allows better utilisation of all available speech and text data through continued self-supervised pre-training and semi-supervised training. We benchmark our approach on Scottish Gaelic, achieving WER reductions of 32% relative over our best fine-tuned Whisper model.

* Accepted to Interspeech 2025

Via

Access Paper or Ask Questions

Evaluating and Adapting Large Language Models to Represent Folktales in Low-Resource Languages

Nov 08, 2024

JA Meaney, Beatrice Alex, William Lamb

Figure 1 for Evaluating and Adapting Large Language Models to Represent Folktales in Low-Resource Languages

Figure 2 for Evaluating and Adapting Large Language Models to Represent Folktales in Low-Resource Languages

Figure 3 for Evaluating and Adapting Large Language Models to Represent Folktales in Low-Resource Languages

Figure 4 for Evaluating and Adapting Large Language Models to Represent Folktales in Low-Resource Languages

Abstract:Folktales are a rich resource of knowledge about the society and culture of a civilisation. Digital folklore research aims to use automated techniques to better understand these folktales, and it relies on abstract representations of the textual data. Although a number of large language models (LLMs) claim to be able to represent low-resource langauges such as Irish and Gaelic, we present two classification tasks to explore how useful these representations are, and three adaptations to improve the performance of these models. We find that adapting the models to work with longer sequences, and continuing pre-training on the domain of folktales improves classification performance, although these findings are tempered by the impressive performance of a baseline SVM with non-contextual features.

Via

Access Paper or Ask Questions