datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
odia-text-corpus
Odia Text Corpus
Dataset Description
This is a comprehensive Odia language text corpus designed for training language models, text generation, and various NLP tasks in Odia (ଓଡ଼ିଆ). The dataset contains high-quality Odia text from multiple sources, providing a rich foundation for Odia language AI development.
Dataset Summary
Language: Odia (ଓଡ଼ିଆ)
Total Records: 649,120
Text Format: Plain Odia text
License: CC-BY-4.0
Use Cases: Language modeling, text… See the full description on the dataset page: https://huggingface.co/datasets/abhilash88/odia-text-corpus.fine_web2_odia_ptodia-instruction-dataset
Odia Instruction Following Dataset
Dataset Description
This is a comprehensive Odia language instruction-following dataset designed for training conversational AI models, chatbots, and instruction-following systems in Odia (ଓଡ଼ିଆ). The dataset contains high-quality instruction-response pairs that enable models to understand and follow instructions in the Odia language.
Dataset Summary
Language: Odia (ଓଡ଼ିଆ)
Total Records: 324,560
Format:… See the full description on the dataset page: https://huggingface.co/datasets/abhilash88/odia-instruction-dataset.odia-eval-benchmark
Odia Eval Benchmark
Dataset Summary
odia-eval-benchmark is a comprehensive, curated evaluation benchmark for Odia (Oriya) natural language understanding and generation. It consolidates 34 publicly available Odia datasets into a single, normalized format covering 7 task families and 121,947 evaluation rows.
This benchmark was built from authoritative sources with three major improvements:
Curated sources - 8 datasets were re-pulled from their… See the full description on the dataset page: https://huggingface.co/datasets/MaelisResearch/odia-eval-benchmark.RAG_Evaluation_Datasetodia-hellaswag
Odia HellaSwag
Odia translation of HellaSwag, a commonsense sentence-completion benchmark. Context and gold endings are translated; distractor endings in all_endings remain in English.
Part of OdiaBench — parallel English–Odia benchmark translations for
evaluating Odia-capable language models.
Splits
Split
Rows
test
10,003
train
39,905
validation
10,042
Total rows: 59,950
Schema
Column
Type
Description
id
int64
Pipeline row… See the full description on the dataset page: https://huggingface.co/datasets/tripathysagar/odia-hellaswag.Odia_English_Sentences_Aligned_81k
🌐 Vaqas AI: Odia-English Sentence-Aligned Precision Corpus
Created and Maintained by Vaqas AI Creator & Owner: Vaqas Ahmed
🚀 Why This Dataset?
Standard parallel corpora often consist of large paragraphs that exceed the 4k or 8k context windows typical of LLM training, leading to truncated data and poor model performance.
The Vaqas AI Sentence-Aligned Corpus was engineered to solve this. We transformed the original vaqasai/Odia_English_News dataset by decomposing dense… See the full description on the dataset page: https://huggingface.co/datasets/vaqasai/Odia_English_Sentences_Aligned_81k.pre_train_odia_dataThe present dataset is compiled by using the following datasets:
CultureaX
Licesnse - ODC-By, CC0, Paper
Source - https://huggingface.co/datasets/uonlp/CulturaX/viewer/or?
49M tokens, 2.9M sentences
Collection of different versions of Ocsar (Commom Crawl data) and mC4 dataset (Common Crawl's web crawl corpus). mC4 forms 66% of CulturaX dataset.
IndicQA
License - cc-by-4.0
Source - https://huggingface.co/datasets/ai4bharat/IndicQA/viewer/indicqa.or
0.23M tokens, 15K… See the full description on the dataset page: https://huggingface.co/datasets/OdiaGenAIdata/pre_train_odia_data.audio_tts_description_odiamalayalam-ocr-datawat24_text_to_text_translationodia-german-parallel-corpus-research
Dataset Summary
This dataset is a high-quality, parallel corpus for Odia (Oriya) to German and German to Odia machine translation. It focuses on the news domain, specifically covering National, International, Sports, Trade, and Science & Technology topics.
The dataset contains 3,676 unique parallel sentence pairs, curated through a hybrid approach combining automated web scraping, manual human translation (Gold Standard), and human-corrected machine translation (Silver… See the full description on the dataset page: https://huggingface.co/datasets/abhinandansamal/odia-german-parallel-corpus-research.
