CoolFace
17 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01IRIIS-RESEARCH /Nepali-Text-Corpus Nepali Text Corpus Overview Nepali-Text-Corpus is a comprehensive collection of approximately 6.4 million articles in the Nepali language. This dataset is the largest text dataset on Nepali Language. It encompasses a diverse range of text types, including news articles, blogs, and more, making it an invaluable resource for researchers, developers, and enthusiasts in the fields of Natural Language Processing (NLP) and computational linguistics. Dataset… See the full description on the dataset page: https://huggingface.co/datasets/IRIIS-RESEARCH/Nepali-Text-Corpus.texttext-generation1M<n<10M10 likes969 downloads1y agoHugging Face02kunishou /J-ResearchCorpus J-ResearchCorpus Update: 2024/3/16言語処理学会第30回年次大会(NLP2024)を含む、論文 1,343 本のデータを追加 2024/2/25言語処理学会誌「自然言語処理」のうち CC-BY-4.0 で公開されている論文 360 本のデータを追加 概要 CC-BY-* ライセンスで公開されている日本語論文や学会誌等から抜粋した高品質なテキストのデータセットです。言語モデルの事前学習や RAG 等でご活用下さい。 今後も CC-BY-* ライセンスの日本語論文があれば追加する予定です。 データ説明 filename : 該当データのファイル名 text : 日本語論文から抽出したテキストデータ category : データソース license : ライセンス credit : クレジット データソース・ライセンス テキスト総文字数 : 約 3,900 万文字 data source num records license note… See the full description on the dataset page: https://huggingface.co/datasets/kunishou/J-ResearchCorpus.text1K<n<10K32 likes211 downloads3y agoHugging Face03fliarbi /urban-heat-research-corpus Urban Heat Research Corpus (UHRC) v1.0 What does the world study, invent and report about urban heat? This dataset puts three records of the same problem side by side: 20,422 research papers on urban heat islands and extreme heat in cities (1990–2025) with the claims their abstracts make, 106,458 news articles about heat (2021–2025) coded for 51 subjects, framings and terms, and 4,123 patent families for heat-mitigation technologies (2006–2024) — plus supplementary tables on the… See the full description on the dataset page: https://huggingface.co/datasets/fliarbi/urban-heat-research-corpus.imagetext-classification100K<n<1M0 likes161 downloads7d agoHugging Face04Deep-Research-Team /Pre-Training-Persian-Corpus-Raw-Texts-DatasetDocument Version: 2.0.0 | Last Updated: 02/13/2026 text10M<n<100M0 likes99 downloads7mo agoHugging Face05NorthernTribe-Research /math-conjecture-training-corpus Math Conjecture Training Corpus (v1) Repository: NorthernTribe-Research/math-conjecture-training-corpus Summary Merged training dataset for unsolved-conjecture-oriented math AI training. Included families: conjecture_core competition structured_reasoning formal_proof Rows Per Split train: 458664 validation: 4920 test: 9765 Rows Per Family conjecture_core: 121 formal_proof: 186239 competition: 63093 structured_reasoning: 223896 Policy… See the full description on the dataset page: https://huggingface.co/datasets/NorthernTribe-Research/math-conjecture-training-corpus.text100K<n<1M0 likes49 downloads6mo agoHugging Face06latam-gpt /LatamGPT-Corpus-1.0-researchgated LatamGPT-Corpus-1.0-research 🌐 Language versions: English | Español | Português 🔗 Project links: Official LatamGPT website | Corpus dashboard 🔓 Open portion: the openly released part of this corpus is published separately as LatamGPT-Corpus-1.0. ⚠️ Controlled Access Level (Blue – Research) This repository is not openly available. It holds data destined exclusively for scientific and academic research, and is managed as a restricted repository under… See the full description on the dataset page: https://huggingface.co/datasets/latam-gpt/LatamGPT-Corpus-1.0-research.texttext-generation10M<n<100M1 likes33 downloads11d agoHugging Face07dibao-research /ming-qing-wenji-corpus Ming-Qing Literary Collections Corpus / 明清別集語料庫 / 명청별집어료고 Dataset Description / 數據集說明 / 데이터셋 설명 Summary / 概要 / 요약 English: The Ming-Qing Literary Collections Corpus is a structured digital corpus of 472 literary collections (bieji 別集) from the Ming (明, 1368–1644) and Qing (清, 1644–1912) dynasties. The texts are sourced from the Siku Quanshu (四庫全書) tradition and include poetry, prose, memorials, essays, letters, and other literary genres by scholars, officials… See the full description on the dataset page: https://huggingface.co/datasets/dibao-research/ming-qing-wenji-corpus.tabulartext-generation10K<n<100K1 likes27 downloads6mo agoHugging Face08codin-research /bo-y-te-corpus-rawgatedtexttext-generationn<1K2 likes20 downloads1y agoHugging Face09trentdoney /agent-memory-research-corpus Agent Memory Research Corpus (AMRC) A public, citable dataset for agent memory research and systems. This dataset catalogues papers, systems, benchmarks, and design patterns related to long-term memory in autonomous agents. It is intended to serve as a canonical reference corpus for researchers and practitioners building memory-augmented agents. Dataset Summary Field Value Repository https://huggingface.co/datasets/trentdoney/agent-memory-research-corpus… See the full description on the dataset page: https://huggingface.co/datasets/trentdoney/agent-memory-research-corpus.textn<1K0 likes19 downloads5mo agoHugging Face10dibao-research /wanli-dibao-corpus Wanli Dibao Corpus / 萬曆邸鈔校訂語料庫 Dataset Description Summary English: The Wanli Dibao Corpus is a structured, proofread digital corpus of the Wanli Dichao (萬曆邸鈔), a collection of manuscript copies of official gazettes (dibao 邸報) from the Wanli reign (1573–1620) of the Ming dynasty. The dibao system was the primary channel of official communication in imperial China, transmitting memorials, edicts, personnel appointments, and policy decisions from the capital to… See the full description on the dataset page: https://huggingface.co/datasets/dibao-research/wanli-dibao-corpus.tabulartext-generation10K<n<100K1 likes18 downloads6mo agoHugging Face11316usman /research_clinical_visit_note_summarization_corpus_mtstext1K<n<10K0 likes17 downloads2y agoHugging Face12GENIAC-Team-Ozaki /J-ResearchCorpus_cleanedtext1K<n<10K0 likes13 downloads2y agoHugging Face13soynade-research /kallaama-retrival-eval-corpus Fleurs-Retrieval-Eval Fleurs-Retrieval-Eval is a cross-lingual speech-to-text retrieval evaluation dataset designed to assess French document retrieval from Wolof speech queries. It is derived from the test split of SIB-Fleurs, a multilingual spoken language understanding benchmark based on the FLEURS corpus. The dataset is constructed using fully natural speech data to provide a realistic evaluation setting, in contrast to synthetic training corpora. For each Wolof speech sample… See the full description on the dataset page: https://huggingface.co/datasets/soynade-research/kallaama-retrival-eval-corpus.textn<1K0 likes10 downloads8mo agoHugging Face14auren-research /privacy-base-corpustext1M<n<10M0 likes4 downloads4mo agoHugging Face15NorthernTribe-Research /maasai-translation-corpusgated Maasai-English Translation Corpus Parallel English↔Maasai translation pairs for low-resource MT, language preservation, and culturally grounded tooling. Overview Total pairs: 9,910 Splits: 8,434 train / 738 valid / 738 test Directions: 4,955 en→mas and 4,955 mas→en Quality tiers: 8,444 gold and 1,466 silver Main sources: 8,444 Bible-derived pairs, 680 cultural manual pairs, 70 knowledge-driven cultural pairs, 132 public-domain Hollis proverb pairs, 504 public-domain… See the full description on the dataset page: https://huggingface.co/datasets/NorthernTribe-Research/maasai-translation-corpus.tabulartranslation10K<n<100K0 likes3 downloads6mo agoHugging Face16codin-research /benh-hoc-corpus-rawgatedtextn<1K0 likes2 downloads1y agoHugging Face17abhinandansamal /odia-german-parallel-corpus-researchgated Dataset Summary This dataset is a high-quality, parallel corpus for Odia (Oriya) to German and German to Odia machine translation. It focuses on the news domain, specifically covering National, International, Sports, Trade, and Science & Technology topics. The dataset contains 3,676 unique parallel sentence pairs, curated through a hybrid approach combining automated web scraping, manual human translation (Gold Standard), and human-corrected machine translation (Silver… See the full description on the dataset page: https://huggingface.co/datasets/abhinandansamal/odia-german-parallel-corpus-research.tabulartranslation1K<n<10K0 likes2 downloads9mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.