CoolFace
12 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01abhilash88 /odia-text-corpus Odia Text Corpus Dataset Description This is a comprehensive Odia language text corpus designed for training language models, text generation, and various NLP tasks in Odia (ଓଡ଼ିଆ). The dataset contains high-quality Odia text from multiple sources, providing a rich foundation for Odia language AI development. Dataset Summary Language: Odia (ଓଡ଼ିଆ) Total Records: 649,120 Text Format: Plain Odia text License: CC-BY-4.0 Use Cases: Language modeling, text… See the full description on the dataset page: https://huggingface.co/datasets/abhilash88/odia-text-corpus.tabulartext-generation100K<n<1M1 likes93 downloads1y agoHugging Face02OdiaGenAIdata /fine_web2_odia_pttabular1M<n<10M0 likes91 downloads2y agoHugging Face03abhilash88 /odia-instruction-dataset Odia Instruction Following Dataset Dataset Description This is a comprehensive Odia language instruction-following dataset designed for training conversational AI models, chatbots, and instruction-following systems in Odia (ଓଡ଼ିଆ). The dataset contains high-quality instruction-response pairs that enable models to understand and follow instructions in the Odia language. Dataset Summary Language: Odia (ଓଡ଼ିଆ) Total Records: 324,560 Format:… See the full description on the dataset page: https://huggingface.co/datasets/abhilash88/odia-instruction-dataset.tabulartext-generation100K<n<1M1 likes69 downloads1y agoHugging Face04MaelisResearch /odia-eval-benchmark Odia Eval Benchmark Dataset Summary odia-eval-benchmark is a comprehensive, curated evaluation benchmark for Odia (Oriya) natural language understanding and generation. It consolidates 34 publicly available Odia datasets into a single, normalized format covering 7 task families and 121,947 evaluation rows. This benchmark was built from authoritative sources with three major improvements: Curated sources - 8 datasets were re-pulled from their… See the full description on the dataset page: https://huggingface.co/datasets/MaelisResearch/odia-eval-benchmark.tabularquestion-answering100K<n<1M2 likes57 downloads1mo agoHugging Face05OdiaGenAI /RAG_Evaluation_Datasettabular1K<n<10K0 likes51 downloads3y agoHugging Face06tripathysagar /odia-hellaswag Odia HellaSwag Odia translation of HellaSwag, a commonsense sentence-completion benchmark. Context and gold endings are translated; distractor endings in all_endings remain in English. Part of OdiaBench — parallel English–Odia benchmark translations for evaluating Odia-capable language models. Splits Split Rows test 10,003 train 39,905 validation 10,042 Total rows: 59,950 Schema Column Type Description id int64 Pipeline row… See the full description on the dataset page: https://huggingface.co/datasets/tripathysagar/odia-hellaswag.tabularmultiple-choice10K<n<100K0 likes26 downloads4mo agoHugging Face07vaqasai /Odia_English_Sentences_Aligned_81k 🌐 Vaqas AI: Odia-English Sentence-Aligned Precision Corpus Created and Maintained by Vaqas AI Creator & Owner: Vaqas Ahmed 🚀 Why This Dataset? Standard parallel corpora often consist of large paragraphs that exceed the 4k or 8k context windows typical of LLM training, leading to truncated data and poor model performance. The Vaqas AI Sentence-Aligned Corpus was engineered to solve this. We transformed the original vaqasai/Odia_English_News dataset by decomposing dense… See the full description on the dataset page: https://huggingface.co/datasets/vaqasai/Odia_English_Sentences_Aligned_81k.tabulartranslation10K<n<100K1 likes13 downloads9mo agoHugging Face08OdiaGenAIdata /pre_train_odia_datagatedThe present dataset is compiled by using the following datasets: CultureaX Licesnse - ODC-By, CC0, Paper Source - https://huggingface.co/datasets/uonlp/CulturaX/viewer/or? 49M tokens, 2.9M sentences Collection of different versions of Ocsar (Commom Crawl data) and mC4 dataset (Common Crawl's web crawl corpus). mC4 forms 66% of CulturaX dataset. IndicQA License - cc-by-4.0 Source - https://huggingface.co/datasets/ai4bharat/IndicQA/viewer/indicqa.or 0.23M tokens, 15K… See the full description on the dataset page: https://huggingface.co/datasets/OdiaGenAIdata/pre_train_odia_data.tabular1M<n<10M0 likes8 downloads3y agoHugging Face09SayantanJoker /audio_tts_description_odiatabular10K<n<100K1 likes8 downloads2y agoHugging Face10OdiaGenAIOCR /malayalam-ocr-datatabularn<1K0 likes8 downloads3mo agoHugging Face11OdiaGenAIdata /wat24_text_to_text_translationtabular100K<n<1M0 likes3 downloads2y agoHugging Face12abhinandansamal /odia-german-parallel-corpus-researchgated Dataset Summary This dataset is a high-quality, parallel corpus for Odia (Oriya) to German and German to Odia machine translation. It focuses on the news domain, specifically covering National, International, Sports, Trade, and Science & Technology topics. The dataset contains 3,676 unique parallel sentence pairs, curated through a hybrid approach combining automated web scraping, manual human translation (Gold Standard), and human-corrected machine translation (Silver… See the full description on the dataset page: https://huggingface.co/datasets/abhinandansamal/odia-german-parallel-corpus-research.tabulartranslation1K<n<10K0 likes2 downloads9mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.