datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Odia-Web-Corpus-v1
Odia Web Corpus v1
The inaugural release of a curated Odia (Oriya) web text corpus. Scraped and filtered from publicly accessible web sources to support Odia language modeling and NLP research.
Dataset Details
Language: Odia (Oriya, ISO 639-3: ory)
Format: JSONL (one JSON object per line)
Size: ~650K documents, ~0.9 GB text
License: CC-BY-4.0
Data Fields
Field
Type
Description
text
string
Cleaned document body
title
string… See the full description on the dataset page: https://huggingface.co/datasets/saidutta69/Odia-Web-Corpus-v1.dolly-odia-15k
Dataset Card for Dolly-Odia-15K
Dataset Summary
This dataset is the Odia-translated version of the Dolly 15K instruction set. In this dataset both English and Odia instruction, input, and output strings are available.
Supported Tasks and Leaderboards
Large Language Model (LLM)
Languages
Odia
Dataset Structure
JSON
Data Fields
instruction (string)
english_instruction (string)
input (string)
english_input (string)
output… See the full description on the dataset page: https://huggingface.co/datasets/OdiaGenAI/dolly-odia-15k.all_combined_bengali_252k
Dataset Card for all_combined_bengali_252K
Dataset Summary
This dataset is a mix of Bengali instruction sets translated from open-source instruction sets:
Dolly,
Alpaca,
ChatDoctor,
Roleplay
GSM
In this dataset Bengali instruction, input, and output strings are available.
Supported Tasks and Leaderboards
Large Language Model (LLM)
Languages
Bengali
Dataset Structure
JSON
Data Fields
output (string)
data_source (string)… See the full description on the dataset page: https://huggingface.co/datasets/OdiaGenAI/all_combined_bengali_252k.gpt-teacher-roleplay-odia-3k
Dataset Card for GPT-Teacher-RolePlay-Odia-3K
Dataset Summary
This dataset is the Odia-translated version of the GPT-Teacher-RolePlay 3K instruction set. In this dataset both English and Odia instruction, input, and output strings are available.
Supported Tasks and Leaderboards
Large Language Model (LLM)
Languages
Odia
Dataset Structure
JSON
Data Fields
instruction (string)
english_instruction (string)
input (string)… See the full description on the dataset page: https://huggingface.co/datasets/OdiaGenAI/gpt-teacher-roleplay-odia-3k.gpt-teacher-instruct-odia-18k
Dataset Card for Odia_GPT-Teacher-Instruct-Odia-18K
Dataset Summary
This dataset is the Odia-translated version of the GPT-Teacher 18K instruction set. In this dataset both English and Odia instruction, input, and output strings are available.
Supported Tasks and Leaderboards
Large Language Model (LLM)
Languages
Odia
Dataset Structure
JSON
Data Fields
instruction (string)
english_instruction (string)
input (string)… See the full description on the dataset page: https://huggingface.co/datasets/OdiaGenAI/gpt-teacher-instruct-odia-18k.odia_domain_context_train_v1
Dataset Card for odia_domain_context_train_v1
Dataset Summary
This dataset contains 10K instructions that span various facets of Odisha's unique identity.
The instructions cover a wide array of subjects, ranging from the culinary delights in 'RECIPES,' the historical significance of 'HISTORICAL PLACES,' and
'TEMPLES OF ODISHA,' to the intellectual pursuits in 'ARITHMETIC,' 'HEALTH,' and 'GEOGRAPHY.'
It also explores the artistic tapestry of Odisha through 'ART AND… See the full description on the dataset page: https://huggingface.co/datasets/OdiaGenAI/odia_domain_context_train_v1.all_combined_odia_171k
Dataset Card for all_combined_odia_171K
Dataset Summary
This dataset is a mix of Odia instruction sets translated from open-source instruction sets.
The Odia instruction sets used are:
dolly-odia-15k
OdiEnCorp_translation_instructions_25k
gpt-teacher-roleplay-odia-3k
Odia_Alpaca_instructions_52k
hardcode_odia_qa_105
In this dataset Odia instruction, input, and output strings are available.
Supported Tasks and Leaderboards
Large Language Model (LLM)… See the full description on the dataset page: https://huggingface.co/datasets/OdiaGenAI/all_combined_odia_171k.OdiEnCorp_translation_instructions_25k
Dataset Card for OdiEnCorp_translation_instructions_25k
Dataset Summary
This dataset is the English-to-Odia translation instruction set. The instruction set is built using the OdienCorp_1.0 English-Odia parallel dataset. The instruction set contains input, and output strings.
Supported Tasks and Leaderboards
Large Language Model (LLM)
Languages
Odia
Dataset Structure
JSON
Data Fields
instruction (string)
output (string)… See the full description on the dataset page: https://huggingface.co/datasets/OdiaGenAI/OdiEnCorp_translation_instructions_25k.odia_context_qa_98k
Dataset Card for odia-qa-98K
Dataset Summary
Supported Tasks and Leaderboards
Large Language Model (LLM)
Languages
Odia
Dataset Structure
JSON
Data Fields
instruction (string)
english_instruction (string)
input (string)
english_input (string)
output (string)
english_output (string)
Licensing Information
This work is licensed under a
Creative Commons Attribution-NonCommercial-ShareAlike 4.0 International License.… See the full description on the dataset page: https://huggingface.co/datasets/OdiaGenAI/odia_context_qa_98k.odia_domain_context_train_v1
Dataset Card for odia_domain_context_train_v1
Dataset Summary
This dataset contains 10K instructions that span various facets of Odisha's unique identity.
The instructions cover a wide array of subjects, ranging from the culinary delights in 'RECIPES,' the historical significance of 'HISTORICAL PLACES,' and
'TEMPLES OF ODISHA,' to the intellectual pursuits in 'ARITHMETIC,' 'HEALTH,' and 'GEOGRAPHY.'
It also explores the artistic tapestry of Odisha through 'ART AND… See the full description on the dataset page: https://huggingface.co/datasets/sarthakprassidh/odia_domain_context_train_v1.
