datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Odia-Web-Corpus-v1
Odia Web Corpus v1
The inaugural release of a curated Odia (Oriya) web text corpus. Scraped and filtered from publicly accessible web sources to support Odia language modeling and NLP research.
Dataset Details
Language: Odia (Oriya, ISO 639-3: ory)
Format: JSONL (one JSON object per line)
Size: ~650K documents, ~0.9 GB text
License: CC-BY-4.0
Data Fields
Field
Type
Description
text
string
Cleaned document body
title
string… See the full description on the dataset page: https://huggingface.co/datasets/saidutta69/Odia-Web-Corpus-v1.odia-handwritten-ocr
Odia Handwritten OCR Dataset
Dataset Description
This dataset contains 182,152 handwritten Odia character images prepared for training OCR models. The dataset covers all 47 OHCS (Odia Handwritten Character Set) characters with balanced class distribution.
Dataset Summary
Language: Odia (ଓଡ଼ିଆ)
Task: Optical Character Recognition (OCR)
Total Images: 182,152
Character Classes: 47
Image Format: Grayscale JPG (32x32 pixels)
Splits: Train (145,717), Validation (18… See the full description on the dataset page: https://huggingface.co/datasets/tell2jyoti/odia-handwritten-ocr.sentiment_analysis_hindiConventions followed to decide the polarity: -
labels consisting of a single value are left undisturbed, i.e. if label = 'pos', then it'll be pos
labels consisting of multiple values separated by '&' are processed. If all the labels are the same ('pos&pos&pos' or 'neg&neg'), then the shortened form of the multiple label is assigned as the final label. For example, if label = 'pos&pos&pos', then final label will be 'pos'.
labels consisting of mixed values ('pos&neg&pos' or 'neg&neu&pos') are… See the full description on the dataset page: https://huggingface.co/datasets/OdiaGenAI/sentiment_analysis_hindi.instruction_set_hindi_1035The dataset has been created using OliveFarm web application.
Following domains have been covered in this dataset:-
Art
Sports (Cricket, Football, Olympics)
Politics
History
Cooking
Environment
Music
Contributors: -
Shahid
Parul.
odia-minimind-dataestsdolly-odia-15k
Dataset Card for Dolly-Odia-15K
Dataset Summary
This dataset is the Odia-translated version of the Dolly 15K instruction set. In this dataset both English and Odia instruction, input, and output strings are available.
Supported Tasks and Leaderboards
Large Language Model (LLM)
Languages
Odia
Dataset Structure
JSON
Data Fields
instruction (string)
english_instruction (string)
input (string)
english_input (string)
output… See the full description on the dataset page: https://huggingface.co/datasets/OdiaGenAI/dolly-odia-15k.all_combined_bengali_252k
Dataset Card for all_combined_bengali_252K
Dataset Summary
This dataset is a mix of Bengali instruction sets translated from open-source instruction sets:
Dolly,
Alpaca,
ChatDoctor,
Roleplay
GSM
In this dataset Bengali instruction, input, and output strings are available.
Supported Tasks and Leaderboards
Large Language Model (LLM)
Languages
Bengali
Dataset Structure
JSON
Data Fields
output (string)
data_source (string)… See the full description on the dataset page: https://huggingface.co/datasets/OdiaGenAI/all_combined_bengali_252k.gpt-teacher-roleplay-odia-3k
Dataset Card for GPT-Teacher-RolePlay-Odia-3K
Dataset Summary
This dataset is the Odia-translated version of the GPT-Teacher-RolePlay 3K instruction set. In this dataset both English and Odia instruction, input, and output strings are available.
Supported Tasks and Leaderboards
Large Language Model (LLM)
Languages
Odia
Dataset Structure
JSON
Data Fields
instruction (string)
english_instruction (string)
input (string)… See the full description on the dataset page: https://huggingface.co/datasets/OdiaGenAI/gpt-teacher-roleplay-odia-3k.roleplay_odiaThe following dataset has been created using camel-ai, by passing various combinations of user and assistant. The dataset was translated to Odia using OdiaGenAI English=>Indic translation app.
odia_domain_context_train_v1
Dataset Card for odia_domain_context_train_v1
Dataset Summary
This dataset contains 10K instructions that span various facets of Odisha's unique identity.
The instructions cover a wide array of subjects, ranging from the culinary delights in 'RECIPES,' the historical significance of 'HISTORICAL PLACES,' and
'TEMPLES OF ODISHA,' to the intellectual pursuits in 'ARITHMETIC,' 'HEALTH,' and 'GEOGRAPHY.'
It also explores the artistic tapestry of Odisha through 'ART AND… See the full description on the dataset page: https://huggingface.co/datasets/OdiaGenAI/odia_domain_context_train_v1.roleplay_hindiThe following dataset has been created using camel-ai, by passing various combinations of user and assistant. The dataset was translated to Hindi using OdiaGenAI English=>Indic translation app.
health_hindi_200Contributors: -
Sonal Khosla
AISSEE_2021_Odia
AISSEE 2021 Odia Dataset
This dataset contains 94 questions from the All-India Sainik School Entrance Exam (class VI) that have been manually verified along with answers. The subjects in the dataset are Math, General Knowledge and Logical Reasoning.
gpt-teacher-instruct-odia-18k
Dataset Card for Odia_GPT-Teacher-Instruct-Odia-18K
Dataset Summary
This dataset is the Odia-translated version of the GPT-Teacher 18K instruction set. In this dataset both English and Odia instruction, input, and output strings are available.
Supported Tasks and Leaderboards
Large Language Model (LLM)
Languages
Odia
Dataset Structure
JSON
Data Fields
instruction (string)
english_instruction (string)
input (string)… See the full description on the dataset page: https://huggingface.co/datasets/OdiaGenAI/gpt-teacher-instruct-odia-18k.roleplay_englishall_combined_odia_171k
Dataset Card for all_combined_odia_171K
Dataset Summary
This dataset is a mix of Odia instruction sets translated from open-source instruction sets.
The Odia instruction sets used are:
dolly-odia-15k
OdiEnCorp_translation_instructions_25k
gpt-teacher-roleplay-odia-3k
Odia_Alpaca_instructions_52k
hardcode_odia_qa_105
In this dataset Odia instruction, input, and output strings are available.
Supported Tasks and Leaderboards
Large Language Model (LLM)… See the full description on the dataset page: https://huggingface.co/datasets/OdiaGenAI/all_combined_odia_171k.Odia_Alpaca_instructions_52k
Dataset Card for Odia_Alpaca_Instruction_52K
Dataset Summary
This dataset is the Odia-translated version of Alpaca 52K instruction set. In this dataset both English and Odia instruction, input, and output strings are available.
Supported Tasks and Leaderboards
Large Language Model (LLM)
Languages
Odia
Dataset Structure
JSON
Data Fields
instruction (string)
english_instruction (string)
input (string)
english_input (string)… See the full description on the dataset page: https://huggingface.co/datasets/OdiaGenAI/Odia_Alpaca_instructions_52k.OdiEnCorp_translation_instructions_25k
Dataset Card for OdiEnCorp_translation_instructions_25k
Dataset Summary
This dataset is the English-to-Odia translation instruction set. The instruction set is built using the OdienCorp_1.0 English-Odia parallel dataset. The instruction set contains input, and output strings.
Supported Tasks and Leaderboards
Large Language Model (LLM)
Languages
Odia
Dataset Structure
JSON
Data Fields
instruction (string)
output (string)… See the full description on the dataset page: https://huggingface.co/datasets/OdiaGenAI/OdiEnCorp_translation_instructions_25k.odia_context_qa_98k
Dataset Card for odia-qa-98K
Dataset Summary
Supported Tasks and Leaderboards
Large Language Model (LLM)
Languages
Odia
Dataset Structure
JSON
Data Fields
instruction (string)
english_instruction (string)
input (string)
english_input (string)
output (string)
english_output (string)
Licensing Information
This work is licensed under a
Creative Commons Attribution-NonCommercial-ShareAlike 4.0 International License.… See the full description on the dataset page: https://huggingface.co/datasets/OdiaGenAI/odia_context_qa_98k.odia-emergency-icu-triage-maths-reasoning-v1odia-emergency-icu-triage-maths-reasoning-v1.11odia_reasoningpre_train_odia_dataThe present dataset is compiled by using the following datasets:
CultureaX
Licesnse - ODC-By, CC0, Paper
Source - https://huggingface.co/datasets/uonlp/CulturaX/viewer/or?
49M tokens, 2.9M sentences
Collection of different versions of Ocsar (Commom Crawl data) and mC4 dataset (Common Crawl's web crawl corpus). mC4 forms 66% of CulturaX dataset.
IndicQA
License - cc-by-4.0
Source - https://huggingface.co/datasets/ai4bharat/IndicQA/viewer/indicqa.or
0.23M tokens, 15K… See the full description on the dataset page: https://huggingface.co/datasets/OdiaGenAIdata/pre_train_odia_data.GPTeacher-Odiaodia-domain-knowledgepre_train_odia_data_processed
About
This dataset is curated from different open-source datasets and prepared Odia data using different techniques (web scraping, OCR) and manually corrected by the Odia native speakers.
The dataset is uniformly processed and contains duplicated entries which can be processed based on usage.
For more details about the data, go through the blog post.
Use Cases
The dataset has many use cases such as:
Pre-training Odia LLM,
Building the Odia BERT model,
Building Odia… See the full description on the dataset page: https://huggingface.co/datasets/OdiaGenAIdata/pre_train_odia_data_processed.odia_domain_context_train_v1
Dataset Card for odia_domain_context_train_v1
Dataset Summary
This dataset contains 10K instructions that span various facets of Odisha's unique identity.
The instructions cover a wide array of subjects, ranging from the culinary delights in 'RECIPES,' the historical significance of 'HISTORICAL PLACES,' and
'TEMPLES OF ODISHA,' to the intellectual pursuits in 'ARITHMETIC,' 'HEALTH,' and 'GEOGRAPHY.'
It also explores the artistic tapestry of Odisha through 'ART AND… See the full description on the dataset page: https://huggingface.co/datasets/sarthakprassidh/odia_domain_context_train_v1.odia-emergency-neuro-sepsis-maths-reasoning-v2odia-critical-care-trauma-fluid-maths-reasoning-v3odia_domain_knowledge_503
