datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
pre_train_odia_dataThe present dataset is compiled by using the following datasets:
CultureaX
Licesnse - ODC-By, CC0, Paper
Source - https://huggingface.co/datasets/uonlp/CulturaX/viewer/or?
49M tokens, 2.9M sentences
Collection of different versions of Ocsar (Commom Crawl data) and mC4 dataset (Common Crawl's web crawl corpus). mC4 forms 66% of CulturaX dataset.
IndicQA
License - cc-by-4.0
Source - https://huggingface.co/datasets/ai4bharat/IndicQA/viewer/indicqa.or
0.23M tokens, 15K… See the full description on the dataset page: https://huggingface.co/datasets/OdiaGenAIdata/pre_train_odia_data.odia-german-parallel-corpus-research
Dataset Summary
This dataset is a high-quality, parallel corpus for Odia (Oriya) to German and German to Odia machine translation. It focuses on the news domain, specifically covering National, International, Sports, Trade, and Science & Technology topics.
The dataset contains 3,676 unique parallel sentence pairs, curated through a hybrid approach combining automated web scraping, manual human translation (Gold Standard), and human-corrected machine translation (Silver… See the full description on the dataset page: https://huggingface.co/datasets/abhinandansamal/odia-german-parallel-corpus-research.
