datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
odia-text-corpus
Odia Text Corpus
Dataset Description
This is a comprehensive Odia language text corpus designed for training language models, text generation, and various NLP tasks in Odia (ଓଡ଼ିଆ). The dataset contains high-quality Odia text from multiple sources, providing a rich foundation for Odia language AI development.
Dataset Summary
Language: Odia (ଓଡ଼ିଆ)
Total Records: 649,120
Text Format: Plain Odia text
License: CC-BY-4.0
Use Cases: Language modeling, text… See the full description on the dataset page: https://huggingface.co/datasets/abhilash88/odia-text-corpus.odia-instruction-dataset
Odia Instruction Following Dataset
Dataset Description
This is a comprehensive Odia language instruction-following dataset designed for training conversational AI models, chatbots, and instruction-following systems in Odia (ଓଡ଼ିଆ). The dataset contains high-quality instruction-response pairs that enable models to understand and follow instructions in the Odia language.
Dataset Summary
Language: Odia (ଓଡ଼ିଆ)
Total Records: 324,560
Format:… See the full description on the dataset page: https://huggingface.co/datasets/abhilash88/odia-instruction-dataset.odia-eval-benchmark
Odia Eval Benchmark
Dataset Summary
odia-eval-benchmark is a comprehensive, curated evaluation benchmark for Odia (Oriya) natural language understanding and generation. It consolidates 34 publicly available Odia datasets into a single, normalized format covering 7 task families and 121,947 evaluation rows.
This benchmark was built from authoritative sources with three major improvements:
Curated sources - 8 datasets were re-pulled from their… See the full description on the dataset page: https://huggingface.co/datasets/MaelisResearch/odia-eval-benchmark.
