Abstract
This contribution describes the collection of a large and diverse corpus for speech recognition and similar tools using crowd-sourced donations. We have built a collection platform inspired by Mozilla Common Voice and specialized it to our needs. We discuss the importance of engaging the community and motivating it to contribute, in our case through competitions. Given the incentive and a platform to easily read in large amounts of utterances, we have observed four cases of speakers freely donating over 10 thousand utterances. We have also seen that women are keener to participate in these events throughout all age groups. Manually verifying a large corpus is a monumental task and we attempt to automatically verify parts of the data using tools like Marosijo and the Montreal Forced Aligner. The method proved helpful, especially for detecting invalid utterances and halving the work needed from crowd-sourced verification.
Original language | English |
---|---|
Title of host publication | 2022 Language Resources and Evaluation Conference, LREC 2022 |
Editors | Nicoletta Calzolari, Frederic Bechet, Philippe Blache, Khalid Choukri, Christopher Cieri, Thierry Declerck, Sara Goggi, Hitoshi Isahara, Bente Maegaard, Joseph Mariani, Helene Mazo, Jan Odijk, Stelios Piperidis |
Publisher | European Language Resources Association (ELRA) |
Pages | 2311-2316 |
Number of pages | 6 |
ISBN (Electronic) | 9791095546726 |
Publication status | Published - 2022 |
Event | 13th International Conference on Language Resources and Evaluation Conference, LREC 2022 - Marseille, France Duration: 20 Jun 2022 → 25 Jun 2022 |
Publication series
Name | 2022 Language Resources and Evaluation Conference, LREC 2022 |
---|
Conference
Conference | 13th International Conference on Language Resources and Evaluation Conference, LREC 2022 |
---|---|
Country/Territory | France |
City | Marseille |
Period | 20/06/22 → 25/06/22 |
Bibliographical note
Publisher Copyright:© European Language Resources Association (ELRA), licensed under CC-BY-NC-4.0.
Other keywords
- Crowd Sourcing
- Icelandic
- Speech corpora