Singh et al.
Indic-centric corpus of 1.16 million sentence pairs improves machine translation for Indian languages
The COILD dataset, built entirely from original Indian language sources across eight domains, consistently outperforms English-pivot baselines in both automatic and human evaluations.
Top University
Kshetrimayum Boynao Singh · Nitin Kumar Mishra · Palash Pratim Dutta · Atai Waris Khan · Aparna Kaushik · Avinash Kumar · +19 more
IIT Patna · IIT Delhi · IIT Guwahati · IIIT Delhi · MIT-MAHE
Research Digest··2 min read
16 million human-verified sentence pairs covering 20 Indian language pairs from four language families.
Why this paper
From MIT-MAHE and 5 others
In one line
COILD is a human-translated parallel corpus covering 20 Indian language pairs that improves machine translation performance.
What we could check
- ·No code link found
- ·No weights link found
- ·No dataset link found
- ·No compute details found
- ·No stated limitations found
- ·No benchmark numbers found
Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.
§