datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Odia-Web-Corpus-v4
Odia Web Corpus v4
Fourth-generation Odia corpus featuring both pretraining data (deduplicated, quality-filtered web text) and instruction-tuning data formatted in ChatML. Built by merging and enhancing v1–v3.
Dataset Details
Language: Odia (ISO 639-3: or)
Format: JSONL.GZ (gzip-compressed JSON lines)
License: CC-BY-SA-4.0
Data Composition
Split
Description
Examples
pretrain_train
Pretraining corpus (train)
~900K… See the full description on the dataset page: https://huggingface.co/datasets/saidutta69/Odia-Web-Corpus-v4.odia-instruction-dataset
Odia Instruction Following Dataset
Dataset Description
This is a comprehensive Odia language instruction-following dataset designed for training conversational AI models, chatbots, and instruction-following systems in Odia (ଓଡ଼ିଆ). The dataset contains high-quality instruction-response pairs that enable models to understand and follow instructions in the Odia language.
Dataset Summary
Language: Odia (ଓଡ଼ିଆ)
Total Records: 324,560
Format:… See the full description on the dataset page: https://huggingface.co/datasets/abhilash88/odia-instruction-dataset.odia-eval-benchmark
Odia Eval Benchmark
Dataset Summary
odia-eval-benchmark is a comprehensive, curated evaluation benchmark for Odia (Oriya) natural language understanding and generation. It consolidates 34 publicly available Odia datasets into a single, normalized format covering 7 task families and 121,947 evaluation rows.
This benchmark was built from authoritative sources with three major improvements:
Curated sources - 8 datasets were re-pulled from their… See the full description on the dataset page: https://huggingface.co/datasets/MaelisResearch/odia-eval-benchmark.instruction_set_hindi_1035The dataset has been created using OliveFarm web application.
Following domains have been covered in this dataset:-
Art
Sports (Cricket, Football, Olympics)
Politics
History
Cooking
Environment
Music
Contributors: -
Shahid
Parul.
odia-eval-benchmark
Odia Eval Benchmark
125,243 evaluation samples across 33 datasets. Zero gating. Zero waiting. Just download and eval.
Why This Is The #1 Odia Evaluation Benchmark
Before this dataset, evaluating Odia language models meant hunting down individual repos,
figuring out each one's format, dealing with broken loaders, and keeping track of
what you've already tested. This is the first and only unified Odia eval benchmark.
Factor
Every Other Option
This… See the full description on the dataset page: https://huggingface.co/datasets/saidutta69/odia-eval-benchmark.odia-arc
Odia ARC-Challenge
Odia translation of AI2 ARC-Challenge. Multiple-choice science questions with English choices and gold labels, plus Odia translations of the question stem and formatted answer block.
Part of OdiaBench — parallel English–Odia benchmark translations for
evaluating Odia-capable language models.
Splits
Split
Rows
test
1,172
train
1,119
validation
299
Total rows: 2,590
Schema
Column
Type
Description
id
int64… See the full description on the dataset page: https://huggingface.co/datasets/tripathysagar/odia-arc.roleplay_odiaThe following dataset has been created using camel-ai, by passing various combinations of user and assistant. The dataset was translated to Odia using OdiaGenAI English=>Indic translation app.
AISSEE_2021_Odia
AISSEE 2021 Odia Dataset
This dataset contains 94 questions from the All-India Sainik School Entrance Exam (class VI) that have been manually verified along with answers. The subjects in the dataset are Math, General Knowledge and Logical Reasoning.
health_hindi_200Contributors: -
Sonal Khosla
odia-gsm8k
Odia GSM8K
Odia translation of GSM8K, grade-school math word problems with chain-of-thought reasoning. Odia fields preserve <<calc=result>> markers and the #### N final-answer line.
Part of OdiaBench — parallel English–Odia benchmark translations for
evaluating Odia-capable language models.
Splits
Split
Rows
test
1,319
train
7,473
Total rows: 8,792
Schema
Column
Type
Description
id
int64
Pipeline row index (0-based, sorted)… See the full description on the dataset page: https://huggingface.co/datasets/tripathysagar/odia-gsm8k.roleplay_englishroleplay_hindiThe following dataset has been created using camel-ai, by passing various combinations of user and assistant. The dataset was translated to Hindi using OdiaGenAI English=>Indic translation app.
odia_reasoning_benchmark
Odia Reasoning Benchmark
This benchmark contains reasoning questions in Odia (logical, Mathematical,Arithmetic, Deductive, Critical Thinking) with answers and optional explanations. Ideal for evaluating Odia QA and reasoning models.
Dataset structure
Column
Description
Question
Reasoning question in Odia
Answer
Correct answer (text or number)
Explanation
Optional step-by-step explanation (some blank)
Type Of Question
Category (e.g., Math, Deductive)… See the full description on the dataset page: https://huggingface.co/datasets/Piyushdash94/odia_reasoning_benchmark.odia-winogrande
Odia Winogrande
Odia translation of the Winogrande pronoun / noun-phrase coreference benchmark (winogrande_xl). Each row keeps English question, answer, and candidate options and adds parallel odia_question / odia_answer fields.
Part of OdiaBench — parallel English–Odia benchmark translations for
evaluating Odia-capable language models.
Splits
Split
Rows
test
1,767
train
40,398
validation
1,267
Total rows: 43,432
Schema
Column
Type… See the full description on the dataset page: https://huggingface.co/datasets/tripathysagar/odia-winogrande.odia-truthfulqa-mc
Odia TruthfulQA (multiple-choice)
Odia translation of the TruthfulQA multiple-choice (MC1) split. Questions and gold answers are translated; distractor choices in mc1_choices remain in English.
Part of OdiaBench — parallel English–Odia benchmark translations for
evaluating Odia-capable language models.
Splits
Split
Rows
validation
817
Total rows: 817
Schema
Column
Type
Description
id
int64
Pipeline row index (0-based, sorted)… See the full description on the dataset page: https://huggingface.co/datasets/tripathysagar/odia-truthfulqa-mc.odia-gemma4-style-polish-mix
OdiaEdgeVoice Gemma4 Style Polish Mix
Weighted dataset for improving Odia chat behavior, punctuation, concise answering,
Romanized Odia handling, and refusal behavior.
Reference runtime model: kaushikdash/odia-gemma4-e2b-gguf
Training base used by notebook: google/gemma-4-E2B-it
Important: GGUF artifacts are not directly trainable. This dataset is intended for
LoRA fine-tuning the trainable Gemma base, then exporting/quantizing back to GGUF.
Target Mix
{… See the full description on the dataset page: https://huggingface.co/datasets/kaushikdash/odia-gemma4-style-polish-mix.hindi_alpaca_dolly_67k_formattedodia-truthfulqa
Odia TruthfulQA (generation)
Odia translation of the TruthfulQA generation split. Open-ended truthfulness questions with English and Odia question/answer pairs.
Part of OdiaBench — parallel English–Odia benchmark translations for
evaluating Odia-capable language models.
Splits
Split
Rows
validation
817
Total rows: 817
Schema
Column
Type
Description
id
int64
Pipeline row index (0-based, sorted)
question
string
English question /… See the full description on the dataset page: https://huggingface.co/datasets/tripathysagar/odia-truthfulqa.Reasoning_ODThe following dataset has been translated from an English reasoning dataset to Odia language to fine-tune and create a Odia reasoning mmodel.
