datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Odia-Web-Corpus-v5
Odia Web Corpus v5
The largest and cleanest Odia text corpus to date — 4.16 million deduplicated documents, 7.74 GB. Built by merging and thoroughly cleaning four source collections.
Dataset Details
Language: Odia (ISO 639-3: or)
Format: 28 sharded Parquet files
Total Size: 7.74 GB
Total Documents: 4,162,804
License: CC-BY-SA-4.0
Cleaning Pipeline
Stage
Removed
Description
Deduplication
30.2%
Exact MD5 hash match
Short lines… See the full description on the dataset page: https://huggingface.co/datasets/saidutta69/Odia-Web-Corpus-v5.odia_pretrain_dataset_v2
Odia Pretrain Dataset v2
12.3 million clean Odia documents. 30+ sources. Zero noise. Zero gating. Just data that actually works for pretraining.
The Origin Story (a.k.a. The Time We Filtered 98% of a Clean Dataset and Learned Our Lesson)
v1 was built from spite. v2 was built from more data.
We took v1 (25+ sources, 10.3M rows) as the foundation and asked: what else can we add?
monsoon-nlp/odia-cleaned -- fully deduplicated against v1 (0 new rows)… See the full description on the dataset page: https://huggingface.co/datasets/saidutta69/odia_pretrain_dataset_v2.Odia-Web-Corpus-v2
Odia Web Corpus v2
Expanded and refined version of the Odia web corpus, converted to Parquet format with dedicated train/test/validation splits for reproducible NLP experimentation.
Dataset Details
Language: Odia (ISO 639-3: or)
Format: Parquet (columnar, compressed)
Splits: train (900K), test (50K), validation (50K)
License: CC-BY-4.0
Data Fields
Field
Type
Description
text
string
Cleaned document text
Usage… See the full description on the dataset page: https://huggingface.co/datasets/saidutta69/Odia-Web-Corpus-v2.Odia-Web-Corpus-v1
Odia Web Corpus v1
The inaugural release of a curated Odia (Oriya) web text corpus. Scraped and filtered from publicly accessible web sources to support Odia language modeling and NLP research.
Dataset Details
Language: Odia (Oriya, ISO 639-3: ory)
Format: JSONL (one JSON object per line)
Size: ~650K documents, ~0.9 GB text
License: CC-BY-4.0
Data Fields
Field
Type
Description
text
string
Cleaned document body
title
string… See the full description on the dataset page: https://huggingface.co/datasets/saidutta69/Odia-Web-Corpus-v1.odia-text-dataset
Odia Text Dataset
This dataset contains Odia text samples.
Usage
from datasets import load_dataset
dataset = load_dataset("hemendra7011/odia-text-dataset")
Dataset Structure
The dataset contains the following fields:
text: The Odia text content
source_file: Source file name
line_number: Line number in source
file_id: File identifier
language: Language (odia)
License
Apache 2.0
Odia-Web-Corpus-v4
Odia Web Corpus v4
Fourth-generation Odia corpus featuring both pretraining data (deduplicated, quality-filtered web text) and instruction-tuning data formatted in ChatML. Built by merging and enhancing v1–v3.
Dataset Details
Language: Odia (ISO 639-3: or)
Format: JSONL.GZ (gzip-compressed JSON lines)
License: CC-BY-SA-4.0
Data Composition
Split
Description
Examples
pretrain_train
Pretraining corpus (train)
~900K… See the full description on the dataset page: https://huggingface.co/datasets/saidutta69/Odia-Web-Corpus-v4.odia-text-corpus
Odia Text Corpus
Dataset Description
This is a comprehensive Odia language text corpus designed for training language models, text generation, and various NLP tasks in Odia (ଓଡ଼ିଆ). The dataset contains high-quality Odia text from multiple sources, providing a rich foundation for Odia language AI development.
Dataset Summary
Language: Odia (ଓଡ଼ିଆ)
Total Records: 649,120
Text Format: Plain Odia text
License: CC-BY-4.0
Use Cases: Language modeling, text… See the full description on the dataset page: https://huggingface.co/datasets/abhilash88/odia-text-corpus.odia_master_data_llama2
Dataset Card for odia_master_data_llama2
Dataset Summary
This dataset is a mix of Odia instruction sets translated from open-source instruction sets and Odia domain knowledge instruction sets.
The Odia instruction sets used are:
odia_domain_context_train_v1
dolly-odia-15k
OdiEnCorp_translation_instructions_25k
gpt-teacher-roleplay-odia-3k
Odia_Alpaca_instructions_52k
hardcode_odia_qa_105
In this dataset Odia instruction, input, and output strings are available.… See the full description on the dataset page: https://huggingface.co/datasets/OdiaGenAI/odia_master_data_llama2.odia-instruction-dataset
Odia Instruction Following Dataset
Dataset Description
This is a comprehensive Odia language instruction-following dataset designed for training conversational AI models, chatbots, and instruction-following systems in Odia (ଓଡ଼ିଆ). The dataset contains high-quality instruction-response pairs that enable models to understand and follow instructions in the Odia language.
Dataset Summary
Language: Odia (ଓଡ଼ିଆ)
Total Records: 324,560
Format:… See the full description on the dataset page: https://huggingface.co/datasets/abhilash88/odia-instruction-dataset.Odia-Web-Corpus-v3
Odia Web Corpus v3
Third iteration of the Odia web corpus with enhanced deduplication, quality filtering, and standardized Parquet splits for pretraining and evaluation.
Dataset Details
Language: Odia (ISO 639-3: or)
Format: Parquet (columnar, compressed)
Splits: train (900K), validation (50K), test (50K)
License: CC-BY-SA-4.0
Data Fields
Field
Type
Description
text
string
Cleaned and filtered document text… See the full description on the dataset page: https://huggingface.co/datasets/saidutta69/Odia-Web-Corpus-v3.odia-eval-benchmark
Odia Eval Benchmark
Dataset Summary
odia-eval-benchmark is a comprehensive, curated evaluation benchmark for Odia (Oriya) natural language understanding and generation. It consolidates 34 publicly available Odia datasets into a single, normalized format covering 7 task families and 121,947 evaluation rows.
This benchmark was built from authoritative sources with three major improvements:
Curated sources - 8 datasets were re-pulled from their… See the full description on the dataset page: https://huggingface.co/datasets/MaelisResearch/odia-eval-benchmark.odia_pretrain_dataset
Odia Pre-training Dataset
10.32 million rows of Odia text. Zero access requests. Zero waiting. Zero gating nonsense.
The Origin Story (a.k.a. How a Pending Access Request Created a Monster)
It all started with a simple request: "Can I please access your Odia pretraining dataset?"
That was months ago. The access request is still pending.
So we did what any reasonable person would do when faced with institutional gatekeeping of a low-resource language's… See the full description on the dataset page: https://huggingface.co/datasets/saidutta69/odia_pretrain_dataset.odia-eval-benchmark
Odia Eval Benchmark
125,243 evaluation samples across 33 datasets. Zero gating. Zero waiting. Just download and eval.
Why This Is The #1 Odia Evaluation Benchmark
Before this dataset, evaluating Odia language models meant hunting down individual repos,
figuring out each one's format, dealing with broken loaders, and keeping track of
what you've already tested. This is the first and only unified Odia eval benchmark.
Factor
Every Other Option
This… See the full description on the dataset page: https://huggingface.co/datasets/saidutta69/odia-eval-benchmark.dolly-odia-15k
Dataset Card for Dolly-Odia-15K
Dataset Summary
This dataset is the Odia-translated version of the Dolly 15K instruction set. In this dataset both English and Odia instruction, input, and output strings are available.
Supported Tasks and Leaderboards
Large Language Model (LLM)
Languages
Odia
Dataset Structure
JSON
Data Fields
instruction (string)
english_instruction (string)
input (string)
english_input (string)
output… See the full description on the dataset page: https://huggingface.co/datasets/OdiaGenAI/dolly-odia-15k.all_combined_bengali_252k
Dataset Card for all_combined_bengali_252K
Dataset Summary
This dataset is a mix of Bengali instruction sets translated from open-source instruction sets:
Dolly,
Alpaca,
ChatDoctor,
Roleplay
GSM
In this dataset Bengali instruction, input, and output strings are available.
Supported Tasks and Leaderboards
Large Language Model (LLM)
Languages
Bengali
Dataset Structure
JSON
Data Fields
output (string)
data_source (string)… See the full description on the dataset page: https://huggingface.co/datasets/OdiaGenAI/all_combined_bengali_252k.gpt-teacher-roleplay-odia-3k
Dataset Card for GPT-Teacher-RolePlay-Odia-3K
Dataset Summary
This dataset is the Odia-translated version of the GPT-Teacher-RolePlay 3K instruction set. In this dataset both English and Odia instruction, input, and output strings are available.
Supported Tasks and Leaderboards
Large Language Model (LLM)
Languages
Odia
Dataset Structure
JSON
Data Fields
instruction (string)
english_instruction (string)
input (string)… See the full description on the dataset page: https://huggingface.co/datasets/OdiaGenAI/gpt-teacher-roleplay-odia-3k.odia_domain_context_train_v1
Dataset Card for odia_domain_context_train_v1
Dataset Summary
This dataset contains 10K instructions that span various facets of Odisha's unique identity.
The instructions cover a wide array of subjects, ranging from the culinary delights in 'RECIPES,' the historical significance of 'HISTORICAL PLACES,' and
'TEMPLES OF ODISHA,' to the intellectual pursuits in 'ARITHMETIC,' 'HEALTH,' and 'GEOGRAPHY.'
It also explores the artistic tapestry of Odisha through 'ART AND… See the full description on the dataset page: https://huggingface.co/datasets/OdiaGenAI/odia_domain_context_train_v1.gpt-teacher-instruct-odia-18k
Dataset Card for Odia_GPT-Teacher-Instruct-Odia-18K
Dataset Summary
This dataset is the Odia-translated version of the GPT-Teacher 18K instruction set. In this dataset both English and Odia instruction, input, and output strings are available.
Supported Tasks and Leaderboards
Large Language Model (LLM)
Languages
Odia
Dataset Structure
JSON
Data Fields
instruction (string)
english_instruction (string)
input (string)… See the full description on the dataset page: https://huggingface.co/datasets/OdiaGenAI/gpt-teacher-instruct-odia-18k.all_combined_odia_171k
Dataset Card for all_combined_odia_171K
Dataset Summary
This dataset is a mix of Odia instruction sets translated from open-source instruction sets.
The Odia instruction sets used are:
dolly-odia-15k
OdiEnCorp_translation_instructions_25k
gpt-teacher-roleplay-odia-3k
Odia_Alpaca_instructions_52k
hardcode_odia_qa_105
In this dataset Odia instruction, input, and output strings are available.
Supported Tasks and Leaderboards
Large Language Model (LLM)… See the full description on the dataset page: https://huggingface.co/datasets/OdiaGenAI/all_combined_odia_171k.odia-gsm8k
Odia GSM8K
Odia translation of GSM8K, grade-school math word problems with chain-of-thought reasoning. Odia fields preserve <<calc=result>> markers and the #### N final-answer line.
Part of OdiaBench — parallel English–Odia benchmark translations for
evaluating Odia-capable language models.
Splits
Split
Rows
test
1,319
train
7,473
Total rows: 8,792
Schema
Column
Type
Description
id
int64
Pipeline row index (0-based, sorted)… See the full description on the dataset page: https://huggingface.co/datasets/tripathysagar/odia-gsm8k.OdiEnCorp_translation_instructions_25k
Dataset Card for OdiEnCorp_translation_instructions_25k
Dataset Summary
This dataset is the English-to-Odia translation instruction set. The instruction set is built using the OdienCorp_1.0 English-Odia parallel dataset. The instruction set contains input, and output strings.
Supported Tasks and Leaderboards
Large Language Model (LLM)
Languages
Odia
Dataset Structure
JSON
Data Fields
instruction (string)
output (string)… See the full description on the dataset page: https://huggingface.co/datasets/OdiaGenAI/OdiEnCorp_translation_instructions_25k.odia_context_10K_llama2_set
Dataset Card for odia_context_10k_llama2_set
Dataset Summary
This dataset contains 10K instructions that span various facets of Odisha's unique identity.
The instructions cover a wide array of subjects, ranging from the culinary delights in 'RECIPES,' the historical significance of 'HISTORICAL PLACES,' and
'TEMPLES OF ODISHA,' to the intellectual pursuits in 'ARITHMETIC,' 'HEALTH,' and 'GEOGRAPHY.'
It also explores the artistic tapestry of Odisha through 'ART AND… See the full description on the dataset page: https://huggingface.co/datasets/OdiaGenAI/odia_context_10K_llama2_set.odia-gemma4-style-polish-mix
OdiaEdgeVoice Gemma4 Style Polish Mix
Weighted dataset for improving Odia chat behavior, punctuation, concise answering,
Romanized Odia handling, and refusal behavior.
Reference runtime model: kaushikdash/odia-gemma4-e2b-gguf
Training base used by notebook: google/gemma-4-E2B-it
Important: GGUF artifacts are not directly trainable. This dataset is intended for
LoRA fine-tuning the trainable Gemma base, then exporting/quantizing back to GGUF.
Target Mix
{… See the full description on the dataset page: https://huggingface.co/datasets/kaushikdash/odia-gemma4-style-polish-mix.odia_context_qa_98k
Dataset Card for odia-qa-98K
Dataset Summary
Supported Tasks and Leaderboards
Large Language Model (LLM)
Languages
Odia
Dataset Structure
JSON
Data Fields
instruction (string)
english_instruction (string)
input (string)
english_input (string)
output (string)
english_output (string)
Licensing Information
This work is licensed under a
Creative Commons Attribution-NonCommercial-ShareAlike 4.0 International License.… See the full description on the dataset page: https://huggingface.co/datasets/OdiaGenAI/odia_context_qa_98k.odia-truthfulqa
Odia TruthfulQA (generation)
Odia translation of the TruthfulQA generation split. Open-ended truthfulness questions with English and Odia question/answer pairs.
Part of OdiaBench — parallel English–Odia benchmark translations for
evaluating Odia-capable language models.
Splits
Split
Rows
validation
817
Total rows: 817
Schema
Column
Type
Description
id
int64
Pipeline row index (0-based, sorted)
question
string
English question /… See the full description on the dataset page: https://huggingface.co/datasets/tripathysagar/odia-truthfulqa.odia_vqa_en_odi_set
Dataset Card for OVQA Instruction Set
Dataset Summary
Odia Visual Question Answering (OVQA) Instruction Set is a multimodal dataset comprising text and images structured in an instruction format, designed for developing Multimodal Large Language Models (MLLMs).
Supported Tasks and Leaderboards
Multimodal Large Language Model (MLLM)
Languages
Odia,
English
Dataset Structure
JSON
Paper
For more details on data preparation… See the full description on the dataset page: https://huggingface.co/datasets/odiagenmllm/odia_vqa_en_odi_set.odia_domain_context_train_v1
Dataset Card for odia_domain_context_train_v1
Dataset Summary
This dataset contains 10K instructions that span various facets of Odisha's unique identity.
The instructions cover a wide array of subjects, ranging from the culinary delights in 'RECIPES,' the historical significance of 'HISTORICAL PLACES,' and
'TEMPLES OF ODISHA,' to the intellectual pursuits in 'ARITHMETIC,' 'HEALTH,' and 'GEOGRAPHY.'
It also explores the artistic tapestry of Odisha through 'ART AND… See the full description on the dataset page: https://huggingface.co/datasets/sarthakprassidh/odia_domain_context_train_v1.
