CoolFace
27 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01saidutta69 /Odia-Web-Corpus-v5 Odia Web Corpus v5 The largest and cleanest Odia text corpus to date — 4.16 million deduplicated documents, 7.74 GB. Built by merging and thoroughly cleaning four source collections. Dataset Details Language: Odia (ISO 639-3: or) Format: 28 sharded Parquet files Total Size: 7.74 GB Total Documents: 4,162,804 License: CC-BY-SA-4.0 Cleaning Pipeline Stage Removed Description Deduplication 30.2% Exact MD5 hash match Short lines… See the full description on the dataset page: https://huggingface.co/datasets/saidutta69/Odia-Web-Corpus-v5.texttext-generation1M<n<10M0 likes287 downloads12d agoHugging Face02saidutta69 /odia_pretrain_dataset_v2 Odia Pretrain Dataset v2 12.3 million clean Odia documents. 30+ sources. Zero noise. Zero gating. Just data that actually works for pretraining. The Origin Story (a.k.a. The Time We Filtered 98% of a Clean Dataset and Learned Our Lesson) v1 was built from spite. v2 was built from more data. We took v1 (25+ sources, 10.3M rows) as the foundation and asked: what else can we add? monsoon-nlp/odia-cleaned -- fully deduplicated against v1 (0 new rows)… See the full description on the dataset page: https://huggingface.co/datasets/saidutta69/odia_pretrain_dataset_v2.texttext-generation10M<n<100M0 likes211 downloads12d agoHugging Face03saidutta69 /Odia-Web-Corpus-v2 Odia Web Corpus v2 Expanded and refined version of the Odia web corpus, converted to Parquet format with dedicated train/test/validation splits for reproducible NLP experimentation. Dataset Details Language: Odia (ISO 639-3: or) Format: Parquet (columnar, compressed) Splits: train (900K), test (50K), validation (50K) License: CC-BY-4.0 Data Fields Field Type Description text string Cleaned document text Usage… See the full description on the dataset page: https://huggingface.co/datasets/saidutta69/Odia-Web-Corpus-v2.texttext-generation1M<n<10M0 likes129 downloads12d agoHugging Face04saidutta69 /Odia-Web-Corpus-v1 Odia Web Corpus v1 The inaugural release of a curated Odia (Oriya) web text corpus. Scraped and filtered from publicly accessible web sources to support Odia language modeling and NLP research. Dataset Details Language: Odia (Oriya, ISO 639-3: ory) Format: JSONL (one JSON object per line) Size: ~650K documents, ~0.9 GB text License: CC-BY-4.0 Data Fields Field Type Description text string Cleaned document body title string… See the full description on the dataset page: https://huggingface.co/datasets/saidutta69/Odia-Web-Corpus-v1.texttext-generation100K<n<1M0 likes126 downloads12d agoHugging Face05hemendra7011 /odia-text-dataset Odia Text Dataset This dataset contains Odia text samples. Usage from datasets import load_dataset dataset = load_dataset("hemendra7011/odia-text-dataset") Dataset Structure The dataset contains the following fields: text: The Odia text content source_file: Source file name line_number: Line number in source file_id: File identifier language: Language (odia) License Apache 2.0 texttext-generation1M<n<10M0 likes113 downloads1y agoHugging Face06saidutta69 /Odia-Web-Corpus-v4 Odia Web Corpus v4 Fourth-generation Odia corpus featuring both pretraining data (deduplicated, quality-filtered web text) and instruction-tuning data formatted in ChatML. Built by merging and enhancing v1–v3. Dataset Details Language: Odia (ISO 639-3: or) Format: JSONL.GZ (gzip-compressed JSON lines) License: CC-BY-SA-4.0 Data Composition Split Description Examples pretrain_train Pretraining corpus (train) ~900K… See the full description on the dataset page: https://huggingface.co/datasets/saidutta69/Odia-Web-Corpus-v4.text-generation1M<n<10M0 likes113 downloads12d agoHugging Face07abhilash88 /odia-text-corpus Odia Text Corpus Dataset Description This is a comprehensive Odia language text corpus designed for training language models, text generation, and various NLP tasks in Odia (ଓଡ଼ିଆ). The dataset contains high-quality Odia text from multiple sources, providing a rich foundation for Odia language AI development. Dataset Summary Language: Odia (ଓଡ଼ିଆ) Total Records: 649,120 Text Format: Plain Odia text License: CC-BY-4.0 Use Cases: Language modeling, text… See the full description on the dataset page: https://huggingface.co/datasets/abhilash88/odia-text-corpus.tabulartext-generation100K<n<1M1 likes93 downloads1y agoHugging Face08OdiaGenAI /odia_master_data_llama2 Dataset Card for odia_master_data_llama2 Dataset Summary This dataset is a mix of Odia instruction sets translated from open-source instruction sets and Odia domain knowledge instruction sets. The Odia instruction sets used are: odia_domain_context_train_v1 dolly-odia-15k OdiEnCorp_translation_instructions_25k gpt-teacher-roleplay-odia-3k Odia_Alpaca_instructions_52k hardcode_odia_qa_105 In this dataset Odia instruction, input, and output strings are available.… See the full description on the dataset page: https://huggingface.co/datasets/OdiaGenAI/odia_master_data_llama2.texttext-generation100K<n<1M1 likes80 downloads3y agoHugging Face09abhilash88 /odia-instruction-dataset Odia Instruction Following Dataset Dataset Description This is a comprehensive Odia language instruction-following dataset designed for training conversational AI models, chatbots, and instruction-following systems in Odia (ଓଡ଼ିଆ). The dataset contains high-quality instruction-response pairs that enable models to understand and follow instructions in the Odia language. Dataset Summary Language: Odia (ଓଡ଼ିଆ) Total Records: 324,560 Format:… See the full description on the dataset page: https://huggingface.co/datasets/abhilash88/odia-instruction-dataset.tabulartext-generation100K<n<1M1 likes69 downloads1y agoHugging Face10saidutta69 /Odia-Web-Corpus-v3 Odia Web Corpus v3 Third iteration of the Odia web corpus with enhanced deduplication, quality filtering, and standardized Parquet splits for pretraining and evaluation. Dataset Details Language: Odia (ISO 639-3: or) Format: Parquet (columnar, compressed) Splits: train (900K), validation (50K), test (50K) License: CC-BY-SA-4.0 Data Fields Field Type Description text string Cleaned and filtered document text… See the full description on the dataset page: https://huggingface.co/datasets/saidutta69/Odia-Web-Corpus-v3.texttext-generation100K<n<1M0 likes65 downloads12d agoHugging Face11MaelisResearch /odia-eval-benchmark Odia Eval Benchmark Dataset Summary odia-eval-benchmark is a comprehensive, curated evaluation benchmark for Odia (Oriya) natural language understanding and generation. It consolidates 34 publicly available Odia datasets into a single, normalized format covering 7 task families and 121,947 evaluation rows. This benchmark was built from authoritative sources with three major improvements: Curated sources - 8 datasets were re-pulled from their… See the full description on the dataset page: https://huggingface.co/datasets/MaelisResearch/odia-eval-benchmark.tabularquestion-answering100K<n<1M2 likes57 downloads1mo agoHugging Face12saidutta69 /odia_pretrain_dataset Odia Pre-training Dataset 10.32 million rows of Odia text. Zero access requests. Zero waiting. Zero gating nonsense. The Origin Story (a.k.a. How a Pending Access Request Created a Monster) It all started with a simple request: "Can I please access your Odia pretraining dataset?" That was months ago. The access request is still pending. So we did what any reasonable person would do when faced with institutional gatekeeping of a low-resource language's… See the full description on the dataset page: https://huggingface.co/datasets/saidutta69/odia_pretrain_dataset.texttext-generation10M<n<100M0 likes38 downloads12d agoHugging Face13saidutta69 /odia-eval-benchmark Odia Eval Benchmark 125,243 evaluation samples across 33 datasets. Zero gating. Zero waiting. Just download and eval. Why This Is The #1 Odia Evaluation Benchmark Before this dataset, evaluating Odia language models meant hunting down individual repos, figuring out each one's format, dealing with broken loaders, and keeping track of what you've already tested. This is the first and only unified Odia eval benchmark. Factor Every Other Option This… See the full description on the dataset page: https://huggingface.co/datasets/saidutta69/odia-eval-benchmark.textmultiple-choice100K<n<1M0 likes36 downloads12d agoHugging Face14OdiaGenAI /dolly-odia-15k Dataset Card for Dolly-Odia-15K Dataset Summary This dataset is the Odia-translated version of the Dolly 15K instruction set. In this dataset both English and Odia instruction, input, and output strings are available. Supported Tasks and Leaderboards Large Language Model (LLM) Languages Odia Dataset Structure JSON Data Fields instruction (string) english_instruction (string) input (string) english_input (string) output… See the full description on the dataset page: https://huggingface.co/datasets/OdiaGenAI/dolly-odia-15k.texttext-generation10K<n<100K0 likes31 downloads3y agoHugging Face15OdiaGenAI /all_combined_bengali_252k Dataset Card for all_combined_bengali_252K Dataset Summary This dataset is a mix of Bengali instruction sets translated from open-source instruction sets: Dolly, Alpaca, ChatDoctor, Roleplay GSM In this dataset Bengali instruction, input, and output strings are available. Supported Tasks and Leaderboards Large Language Model (LLM) Languages Bengali Dataset Structure JSON Data Fields output (string) data_source (string)… See the full description on the dataset page: https://huggingface.co/datasets/OdiaGenAI/all_combined_bengali_252k.texttext-generation100K<n<1M10 likes29 downloads3y agoHugging Face16OdiaGenAI /gpt-teacher-roleplay-odia-3k Dataset Card for GPT-Teacher-RolePlay-Odia-3K Dataset Summary This dataset is the Odia-translated version of the GPT-Teacher-RolePlay 3K instruction set. In this dataset both English and Odia instruction, input, and output strings are available. Supported Tasks and Leaderboards Large Language Model (LLM) Languages Odia Dataset Structure JSON Data Fields instruction (string) english_instruction (string) input (string)… See the full description on the dataset page: https://huggingface.co/datasets/OdiaGenAI/gpt-teacher-roleplay-odia-3k.texttext-generation1K<n<10K5 likes28 downloads3y agoHugging Face17OdiaGenAI /odia_domain_context_train_v1 Dataset Card for odia_domain_context_train_v1 Dataset Summary This dataset contains 10K instructions that span various facets of Odisha's unique identity. The instructions cover a wide array of subjects, ranging from the culinary delights in 'RECIPES,' the historical significance of 'HISTORICAL PLACES,' and 'TEMPLES OF ODISHA,' to the intellectual pursuits in 'ARITHMETIC,' 'HEALTH,' and 'GEOGRAPHY.' It also explores the artistic tapestry of Odisha through 'ART AND… See the full description on the dataset page: https://huggingface.co/datasets/OdiaGenAI/odia_domain_context_train_v1.texttext-generation10K<n<100K0 likes24 downloads3y agoHugging Face18OdiaGenAI /gpt-teacher-instruct-odia-18k Dataset Card for Odia_GPT-Teacher-Instruct-Odia-18K Dataset Summary This dataset is the Odia-translated version of the GPT-Teacher 18K instruction set. In this dataset both English and Odia instruction, input, and output strings are available. Supported Tasks and Leaderboards Large Language Model (LLM) Languages Odia Dataset Structure JSON Data Fields instruction (string) english_instruction (string) input (string)… See the full description on the dataset page: https://huggingface.co/datasets/OdiaGenAI/gpt-teacher-instruct-odia-18k.texttext-generation10K<n<100K1 likes23 downloads3y agoHugging Face19OdiaGenAI /all_combined_odia_171k Dataset Card for all_combined_odia_171K Dataset Summary This dataset is a mix of Odia instruction sets translated from open-source instruction sets. The Odia instruction sets used are: dolly-odia-15k OdiEnCorp_translation_instructions_25k gpt-teacher-roleplay-odia-3k Odia_Alpaca_instructions_52k hardcode_odia_qa_105 In this dataset Odia instruction, input, and output strings are available. Supported Tasks and Leaderboards Large Language Model (LLM)… See the full description on the dataset page: https://huggingface.co/datasets/OdiaGenAI/all_combined_odia_171k.texttext-generation100K<n<1M5 likes22 downloads3y agoHugging Face20tripathysagar /odia-gsm8k Odia GSM8K Odia translation of GSM8K, grade-school math word problems with chain-of-thought reasoning. Odia fields preserve <<calc=result>> markers and the #### N final-answer line. Part of OdiaBench — parallel English–Odia benchmark translations for evaluating Odia-capable language models. Splits Split Rows test 1,319 train 7,473 Total rows: 8,792 Schema Column Type Description id int64 Pipeline row index (0-based, sorted)… See the full description on the dataset page: https://huggingface.co/datasets/tripathysagar/odia-gsm8k.texttext-generation1K<n<10K0 likes22 downloads4mo agoHugging Face21OdiaGenAI /OdiEnCorp_translation_instructions_25k Dataset Card for OdiEnCorp_translation_instructions_25k Dataset Summary This dataset is the English-to-Odia translation instruction set. The instruction set is built using the OdienCorp_1.0 English-Odia parallel dataset. The instruction set contains input, and output strings. Supported Tasks and Leaderboards Large Language Model (LLM) Languages Odia Dataset Structure JSON Data Fields instruction (string) output (string)… See the full description on the dataset page: https://huggingface.co/datasets/OdiaGenAI/OdiEnCorp_translation_instructions_25k.texttext-generation10K<n<100K0 likes16 downloads3y agoHugging Face22OdiaGenAI /odia_context_10K_llama2_set Dataset Card for odia_context_10k_llama2_set Dataset Summary This dataset contains 10K instructions that span various facets of Odisha's unique identity. The instructions cover a wide array of subjects, ranging from the culinary delights in 'RECIPES,' the historical significance of 'HISTORICAL PLACES,' and 'TEMPLES OF ODISHA,' to the intellectual pursuits in 'ARITHMETIC,' 'HEALTH,' and 'GEOGRAPHY.' It also explores the artistic tapestry of Odisha through 'ART AND… See the full description on the dataset page: https://huggingface.co/datasets/OdiaGenAI/odia_context_10K_llama2_set.texttext-generation10K<n<100K1 likes15 downloads3y agoHugging Face23kaushikdash /odia-gemma4-style-polish-mix OdiaEdgeVoice Gemma4 Style Polish Mix Weighted dataset for improving Odia chat behavior, punctuation, concise answering, Romanized Odia handling, and refusal behavior. Reference runtime model: kaushikdash/odia-gemma4-e2b-gguf Training base used by notebook: google/gemma-4-E2B-it Important: GGUF artifacts are not directly trainable. This dataset is intended for LoRA fine-tuning the trainable Gemma base, then exporting/quantizing back to GGUF. Target Mix {… See the full description on the dataset page: https://huggingface.co/datasets/kaushikdash/odia-gemma4-style-polish-mix.texttext-generation10K<n<100K0 likes15 downloads5mo agoHugging Face24OdiaGenAI /odia_context_qa_98k Dataset Card for odia-qa-98K Dataset Summary Supported Tasks and Leaderboards Large Language Model (LLM) Languages Odia Dataset Structure JSON Data Fields instruction (string) english_instruction (string) input (string) english_input (string) output (string) english_output (string) Licensing Information This work is licensed under a Creative Commons Attribution-NonCommercial-ShareAlike 4.0 International License.… See the full description on the dataset page: https://huggingface.co/datasets/OdiaGenAI/odia_context_qa_98k.texttext-generation10K<n<100K0 likes12 downloads3y agoHugging Face25tripathysagar /odia-truthfulqa Odia TruthfulQA (generation) Odia translation of the TruthfulQA generation split. Open-ended truthfulness questions with English and Odia question/answer pairs. Part of OdiaBench — parallel English–Odia benchmark translations for evaluating Odia-capable language models. Splits Split Rows validation 817 Total rows: 817 Schema Column Type Description id int64 Pipeline row index (0-based, sorted) question string English question /… See the full description on the dataset page: https://huggingface.co/datasets/tripathysagar/odia-truthfulqa.texttext-generationn<1K0 likes9 downloads4mo agoHugging Face26odiagenmllm /odia_vqa_en_odi_setgated Dataset Card for OVQA Instruction Set Dataset Summary Odia Visual Question Answering (OVQA) Instruction Set is a multimodal dataset comprising text and images structured in an instruction format, designed for developing Multimodal Large Language Models (MLLMs). Supported Tasks and Leaderboards Multimodal Large Language Model (MLLM) Languages Odia, English Dataset Structure JSON Paper For more details on data preparation… See the full description on the dataset page: https://huggingface.co/datasets/odiagenmllm/odia_vqa_en_odi_set.imagetext-generation10K<n<100K3 likes7 downloads2y agoHugging Face27sarthakprassidh /odia_domain_context_train_v1 Dataset Card for odia_domain_context_train_v1 Dataset Summary This dataset contains 10K instructions that span various facets of Odisha's unique identity. The instructions cover a wide array of subjects, ranging from the culinary delights in 'RECIPES,' the historical significance of 'HISTORICAL PLACES,' and 'TEMPLES OF ODISHA,' to the intellectual pursuits in 'ARITHMETIC,' 'HEALTH,' and 'GEOGRAPHY.' It also explores the artistic tapestry of Odisha through 'ART AND… See the full description on the dataset page: https://huggingface.co/datasets/sarthakprassidh/odia_domain_context_train_v1.texttext-generation10K<n<100K0 likes7 downloads8mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.