Paper ID: 2307.12282

Milimili. Collecting Parallel Data via Crowdsourcing

Alexander Antonov

We present a methodology for gathering a parallel corpus through crowdsourcing, which is more cost-effective than hiring professional translators, albeit at the expense of quality. Additionally, we have made available experimental parallel data collected for Chechen-Russian and Fula-English language pairs.

Submitted: Jul 23, 2023

Topics

Parallel Corpus
Language Pair
Crowdsourcing Context
Parallel Data
Document Translation

Links

arXiv PDF