Kyrgyz is spoken by millions of people, yet it remains significantly underrepresented in modern language technology. While working on English-to-Kyrgyz machine translation, I found that existing systems can produce grammatical errors, lose nuance, or mistranslate less common expressions. I built this project to improve Kyrgyz translation through better data and open collaboration.
This project focuses on improving English-to-Kyrgyz machine translation by building and refining a large bilingual dataset. The dataset contains more than 200,000 English–Kyrgyz sentence pairs that can be used to train and evaluate translation models.
Because Kyrgyz is a low-resource language, high-quality bilingual data is especially valuable. Community corrections can help identify grammatical mistakes, unnatural phrasing, missing meaning, and translation errors that automated systems may overlook.
The dataset contains 200,000+ English–Kyrgyz sentence pairs. It is provided as a 7z archive for researchers, developers, and anyone interested in working with Kyrgyz language technology.
Download the English–Kyrgyz dataset and use it for research, development, or evaluation.
Download 200K+ DatasetFile format: JSONL inside a 7z archive.
If you find a sentence that is incorrect, unnatural, or does not preserve the meaning of the English original, submit a correction. Please include the pair ID so the sentence can be located in the dataset.
Use the form below to report an issue in the dataset. Your correction may be reviewed and incorporated into future versions.
Prefer email? Send corrections to suiunsingapore@gmail.com